Large model inference method, computing device, storage device, and heterogeneous inference system
Patent Information
- Application Number
- CN202510344715.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]然而,存储KV矩阵需占用存储组件中大量的存储空间,当大模型的输入序列较长或者批量处理大模型的输入序列时,存储组件无法提供足够的存储空间,从而限制了大模型推理的吞吐量,导致大模型推理效率较低
Smart Images

Figure CN122797705A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a large-scale model reasoning method, computing device, storage device, and heterogeneous reasoning system. Background Technology
[0002] In artificial intelligence (AI) model inference scenarios, processing-in-memory (PIM) technology can be used to offload some computational tasks that were originally performed by the neural network processing unit (NPU) to storage components with computing capabilities, thereby reducing long-distance data transmission between the NPU and storage components.
[0003] In related technologies, taking generative large models using the Transformer architecture as an example, PIM technology can be used to offload the computational tasks of the attention layer in the Transformer layer to a storage component with computational capabilities. Thus, during the decoding phase of large model inference, when the storage component performs the computational tasks of the attention layer to generate new tokens, it can quickly obtain the KV matrix (i.e., the key and value matrices of the tokens) of the generated tokens, reducing long-distance data transmission.
[0004] However, storing the KV matrix requires a large amount of storage space in the storage component. When the input sequence of the large model is long or the input sequence of the large model is processed in batches, the storage component cannot provide enough storage space, thus limiting the throughput of the large model inference and resulting in low inference efficiency of the large model. Summary of the Invention
[0005] This application provides a large model inference method, computing device, storage device, and heterogeneous inference system, which can improve the throughput and efficiency of large model inference.
[0006] Firstly, a large-model inference method is provided, applied to heterogeneous inference systems. The heterogeneous inference system includes a processor and a storage component. The storage component includes storage units and computing units, which are integrated to enable the storage component to have computing capabilities. "Storage unit integrating computing units" means that computing units are integrated within the storage unit; correspondingly, the storage component is a storage product with in-memory computing capabilities. Alternatively, "storage unit integrating computing units" means that computing units are integrated around or adjacent to the storage unit's physical layer; correspondingly, the storage component is a storage product with near-data processing capabilities. The method includes:
[0007] The processor determines a first execution strategy based on at least one inference request of the large model. The first execution strategy indicates a first part of the computational task that needs to be offloaded to the storage component for execution during the execution of at least one inference request.
[0008] The processor and storage component jointly execute at least one inference request according to a first execution strategy, wherein the computing unit executes a computing task offloaded to the storage component based on inference data stored in the storage unit, the inference data including at least one of model weights and input data of the computing task;
[0009] During the execution of at least one inference request, the processor determines a second execution strategy based on a first execution strategy and the memory usage state of the storage component. The second execution strategy indicates a second part of the computational task that needs to be offloaded to the storage component during the execution of at least one inference request.
[0010] The processor and storage components continue to jointly execute at least one inference request in accordance with the second execution strategy.
[0011] In the above method, the processor determines the execution strategy for at least one inference request to indicate the computational tasks that need to be offloaded to the storage component for execution. The processor and the storage component with computational capabilities then jointly execute the at least one inference request according to the execution strategy. During this process, the processor can sense the memory load of the storage component and adjust the execution strategy in a timely manner based on the memory load. For example, it can schedule a computational task that originally needed to be offloaded to the storage component to be executed on the processor side. This allows the storage component to promptly release the storage space occupied by the data related to that computational task. In this way, by adjusting the execution strategy in a timely manner based on the memory load of the storage component during large model inference, computational resources can be allocated rationally, effectively improving the system's resource utilization and adapting to different inference needs. This allows large models to quickly output inference results even when processing long sequences or batch processing input sequences, thereby improving the throughput of large model inference, i.e., improving the efficiency of large model inference.
[0012] In some embodiments, the computational tasks offloaded to the storage component include at least one of the attention computation task of the Transformer layer in the large model and the feedforward network computation task; wherein the attention computation task includes at least one of the query-key-value QKV matrix generation task, attention score calculation task, normalization calculation task, context vector calculation task, and linear projection task.
[0013] It should be understood that the Transformer layer is a core component of large models. Offloading the computational tasks of the Transformer layer to the storage component allows data to be processed internally, significantly reducing data transfer time and thus improving task execution efficiency. For example, in the attention score calculation task, performing the computation within the storage component avoids transferring a large number of intermediate results to the processor, thereby alleviating the pressure on memory bandwidth. Moreover, offloading both the attention calculation task and the feedforward network calculation task to the storage component enables complete offloading of the Transformer layer, minimizing data transfer time and improving the inference efficiency of large models.
[0014] In some embodiments, the computation unit includes a vector unit and a matrix-vector multiplication unit; the computation unit performs computation tasks offloaded to the storage component based on the inference data stored in the storage unit, including: the vector unit performing a normalization computation task offloaded to the storage component based on the inference data; and the matrix-vector multiplication unit performing tasks other than the normalization computation task among the computation tasks offloaded to the storage component based on the inference data.
[0015] The vector unit provides the computational power for operations between vectors, i.e., vector computational power; the matrix-vector multiplication unit provides the computational power for multiplying matrices and vectors, i.e., general matrix-vector multiplication (GEMV) computational power. This clearly defined division of labor allows for optimization based on the characteristics of different computational tasks, improving computational parallelism and efficiency, further enhancing the performance of the storage components in performing computational tasks, and making large model inference more efficient.
[0016] In some embodiments, the processor determines a first execution strategy based on at least one inference request of the large model, including: the processor determines the first execution strategy based on the number of at least one inference request, the number of storage components, the number of processors, the number of storage units in the storage components, and the number of attention heads in the Transformer layer of the large model; wherein, one storage unit in the storage component corresponds to at least one attention head in the Transformer layer of the large model.
[0017] In this way, the processor can comprehensively determine the first execution strategy based on multiple factors, and the storage unit is associated with the number of attention heads in the Transformer layer. This allows for the reasonable allocation of the correspondence between the storage units of the storage component and the attention heads in the Transformer layer, ensuring the reasonable allocation and effective utilization of computing resources, and improving the overall performance and resource utilization of large model inference.
[0018] In some embodiments, the storage component includes a plurality of storage cell groups, each storage cell group including a plurality of storage cells; the processor and the storage component jointly execute at least one inference request in accordance with a first execution strategy, including: the processor and each storage cell group in the plurality of storage cell groups jointly execute a plurality of inference requests in accordance with the first execution strategy, wherein one storage cell group corresponds to at least one inference request.
[0019] By using the above-described grouping execution method, multiple inference requests can be processed in parallel, making full use of the parallel computing capabilities of the storage components, improving inference efficiency, and accelerating the processing speed of inference requests. This method is particularly suitable for scenarios that require processing a large number of inference requests.
[0020] In some embodiments, during the execution of at least one inference request, the processor determines a second execution strategy based on a first execution strategy and the memory usage state of the storage component, including: during the execution of the decoding phase of at least one inference request, the processor determines a second execution strategy and a memory reclamation strategy based on the first execution strategy and the memory usage state of the storage component, wherein the memory reclamation strategy indicates the inference data that needs to be unloaded from the storage component during the execution of the decoding phase of at least one inference request;
[0021] The method further includes: the processor sending a memory reclamation request to the storage component according to a memory reclamation strategy; and the storage component unloading inference data from the storage component according to the instructions of the memory reclamation request.
[0022] By using the above method, during the decoding phase of the inference request, the second execution strategy and memory reclamation strategy are determined based on the first execution strategy and the memory usage status of the storage component, and the memory reclamation operation is performed. This can release the memory space occupied by inference data that is no longer needed on the storage component in a timely manner, avoid the waste of memory resources, prevent memory overflow and other problems, and ensure that the storage component always has enough memory to execute subsequent computing tasks during the inference process, thereby stabilizing the performance of large model inference.
[0023] In some embodiments, the processor determines a second execution strategy and a memory reclamation strategy based on a first execution strategy and the memory usage state of the storage component, including: the processor determining first reference information based on the first execution strategy and the memory usage state, the first reference information indicating the memory usage of the storage component when continuing to execute at least one inference request according to the first execution strategy; and the processor determining the second execution strategy and the memory reclamation strategy based on the first reference information and the type of the first part of the computation task.
[0024] By determining the first reference information, namely the memory usage of the storage component when continuing to execute according to the first execution strategy, and then determining the second execution strategy and memory reclamation strategy based on this information and the first part of the computing task type, the adjustment of the execution strategy is more based on the actual memory usage and computing task characteristics. This enables more precise optimization of memory usage and computing resource allocation, and improves the system's adaptability and inference efficiency.
[0025] In some embodiments, the first part of the computation task includes an attention computation task and a feedforward network computation task in the Transformer layer of the large model; the processor determines a second execution strategy and a memory reclamation strategy based on the type of computation task indicated by the first reference information and the first execution strategy, including: the processor determines second reference information based on the first reference information and the feedforward network computation task in the first part of the computation task, the second reference information indicating the memory usage of the storage component after unloading the inference data of the feedforward network computation task in the first part of the computation task from the storage component; the processor determines the second execution strategy and the memory reclamation strategy based on the second reference information and the first execution strategy.
[0026] In some embodiments, the processor determines a second execution strategy and a memory reclamation strategy based on second reference information and a first execution strategy, including: if the memory usage indicated by the second reference information meets the conditions, the processor determines the second execution strategy and the memory reclamation strategy based on the second reference information and the first execution strategy; if the memory usage indicated by the second reference information does not meet the conditions, the processor determines third reference information based on the second reference information and the attention computation task in the first part of the computation task, and determines the second execution strategy and the memory reclamation strategy based on the third reference information and the first execution strategy, wherein the third reference information indicates the memory usage of the storage component after unloading the inference data of the feedforward network computation task and the attention computation task in the first part of the computation task from the storage component.
[0027] In this manner, the processor first determines whether the inference data corresponding to the feedforward network computation task needs to be reclaimed. If, after reclaiming the inference data, the memory load of the storage component still does not meet the requirements, it then determines whether the inference data corresponding to at least one of the QKV matrix generation task and the linear projection task needs to be reclaimed. Based on the determined inference data that needs to be reclaimed, a memory reclamation strategy is generated, and the inference data stored on the storage component is unloaded according to the memory reclamation strategy. It should be understood that since the feedforward network computation task is implemented through matrix multiplication operators and is a computationally intensive task, determining whether the inference data corresponding to the feedforward network computation task needs to be reclaimed first can prioritize scheduling computationally intensive tasks to the processor side for execution, thereby improving computational efficiency.
[0028] In some embodiments, the processor includes a central processing unit (CPU) and an accelerator, the CPU being used to control the accelerator, and the CPU and the accelerator being connected via a bus; the processor and the storage component jointly execute at least one inference request according to a first execution strategy, including: the CPU controlling the accelerator to execute tasks in at least one inference request other than computational tasks offloaded to the storage component according to the first execution strategy, and controlling the accelerator to offload computational tasks to the storage component for execution.
[0029] The heterogeneous processor architecture described above fully leverages the control capabilities of the CPU and the computational advantages of the accelerator, rationally allocating computational tasks. Tasks other than those executed by the storage components are offloaded to the accelerator, improving computational efficiency. At the same time, the CPU's control over the accelerator can better coordinate system resources, ensuring the smooth progress of the inference process.
[0030] In some embodiments, the storage component is a high-bandwidth memory (HBM), and the storage component and the processor are connected via a bus; wherein, the storage unit is a storage bank in the HBM, and the bank is integrated with the computing unit.
[0031] By utilizing the high bandwidth advantage of HBM, data can be transmitted quickly, data transmission latency can be reduced, computational efficiency can be improved, and the performance and speed of large model inference can be enhanced.
[0032] Secondly, a computing device for performing large model inference is provided, wherein the computing device and a storage component are connected, the storage component includes a storage unit and a computing unit, the storage unit and the computing unit are integrated, and the computing device includes:
[0033] The first determining module is used to determine a first execution strategy based on at least one inference request of the large model. The first execution strategy indicates a first part of the computational task that needs to be offloaded to the storage component for execution during the execution of at least one inference request.
[0034] An execution module is configured to execute at least one computational task in an inference request that has not been offloaded to a storage component, in accordance with a first execution strategy.
[0035] The second determining module is used to determine a second execution strategy based on the first execution strategy and the memory usage state of the storage component during the execution of at least one inference request. The second execution strategy indicates a second part of the computational task that needs to be offloaded to the storage component for execution during the execution of at least one inference request.
[0036] The execution module is also configured to continue executing at least one computational task in an inference request that has not been offloaded to the storage component, in accordance with the second execution strategy.
[0037] In some embodiments, the computational tasks to be offloaded to the storage component include at least one of the attention computation task of the Transformer layer in the large model and the feedforward network computation task; wherein, the attention computation task includes at least one of the QKV matrix generation task, attention score calculation task, normalization calculation task, context vector calculation task and linear projection task.
[0038] In some embodiments, the first determining module is configured to: determine a first execution strategy based on the number of at least one inference request, the number of storage components, the number of processors, the number of storage units in the storage components, and the number of attention heads in the Transformer layer of the large model; wherein, one storage unit in the storage component corresponds to at least one attention head in the Transformer layer of the large model.
[0039] In some embodiments, the second determining module is configured to: determine a second execution strategy and a memory reclamation strategy based on a first execution strategy and the memory usage state of a storage component during the execution of the decoding phase of at least one inference request; the memory reclamation strategy indicates the inference data that needs to be unloaded from the storage component during the execution of the decoding phase of at least one inference request.
[0040] In some embodiments, the computing device further includes a sending module for: sending a memory reclamation request to the storage component in accordance with a memory reclamation strategy.
[0041] In some embodiments, the second determining module is configured to: determine first reference information based on a first execution strategy and memory usage status, the first reference information indicating the memory usage of the storage component when continuing to execute at least one inference request according to the first execution strategy; and determine a second execution strategy and a memory reclamation strategy based on the first reference information and the type of the first part of the computation task.
[0042] In some embodiments, the first computational task includes an attention computation task and a feedforward network computation task in the Transformer layer of the large model; the second determining module is configured to: determine second reference information based on the first reference information and the feedforward network computation task in the first computational task, the second reference information indicating the memory usage of the storage component after unloading the inference data of the feedforward network computation task in the first computational task from the storage component; and determine a second execution strategy and a memory reclamation strategy based on the second reference information and the first execution strategy.
[0043] In some embodiments, the second determining module is configured to: if the memory usage indicated by the second reference information meets the conditions, determine a second execution strategy and a memory reclamation strategy based on the second reference information and the first execution strategy; if the memory usage indicated by the second reference information does not meet the conditions, determine third reference information based on the second reference information and the attention computing task in the first part of the computing task, and determine the second execution strategy and the memory reclamation strategy based on the third reference information and the first execution strategy, wherein the third reference information indicates the memory usage of the storage component after unloading the inference data of the feedforward network computing task and the attention computing task in the first part of the computing task from the storage component.
[0044] Thirdly, a storage device for performing large model inference is provided, wherein the storage device is connected to a processor, and the storage device includes storage units and computing units, wherein the storage units and computing units are integrated.
[0045] This storage device is used for:
[0046] At least one inference request of a large model is executed according to a first execution strategy. The first execution strategy indicates that a first part of the computation task needs to be offloaded to the storage device for execution during the execution of at least one inference request. The computation unit is used to execute the computation task offloaded to the storage device according to the inference data stored in the storage unit. The inference data includes at least one of model weights and input data of the computation task.
[0047] During the execution of at least one inference request, the execution of at least one inference request continues in accordance with a second execution strategy, which indicates that a second part of the computational task to be offloaded to the storage device during the execution of at least one inference request.
[0048] In some embodiments, the computational tasks offloaded to the storage device for execution include at least one of the attention computation task of the Transformer layer in the large model and the feedforward network computation task; wherein the attention computation task includes at least one of the QKV matrix generation task, attention score computation task, normalization computation task, context vector computation task, and linear projection task.
[0049] In some embodiments, the computation unit includes a vector unit and a matrix-vector multiplication unit;
[0050] Vector units are used to perform normalization computation tasks offloaded to storage devices based on inference data;
[0051] The matrix-vector multiplication unit is used to perform tasks other than normalization calculations from the computational tasks offloaded to the storage device based on inference data.
[0052] In some embodiments, the storage device is further configured to: unload inference data on the storage device according to an instruction from a memory reclamation request sent by the processor.
[0053] In some embodiments, the storage device includes a high-bandwidth memory (HBM), and the HBM and processor are connected via a bus; wherein, the storage unit is a storage bank in the HBM, and the bank is integrated with the computing unit.
[0054] Fourthly, a heterogeneous inference system is provided, which includes a computing device as provided in the second aspect or any possible implementation thereof, and a storage device as provided in the third aspect or any possible implementation thereof.
[0055] Fifthly, a computer-readable storage medium is provided for storing at least one piece of program code for implementing the large model inference method provided by the first aspect or any possible implementation thereof. The storage medium includes, but is not limited to, volatile memory, such as random access memory, and non-volatile memory, such as flash memory, hard disk drive (HDD), and solid-state drive (SSD).
[0056] Sixthly, a computer program product is provided for implementing the large model inference method provided in the first aspect or any possible implementation thereof. The computer program product can be a software installation package, which can be downloaded and executed on a heterogeneous inference system when the aforementioned large model inference method needs to be implemented. Attached Figure Description
[0057] Figure 1 This is a schematic diagram of the Transformer layer structure in a large model;
[0058] Figure 2 This is a schematic diagram of an implementation environment provided in an embodiment of this application;
[0059] Figure 3 This is a schematic diagram of a heterogeneous inference system provided in an embodiment of this application;
[0060] Figure 4 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0061] Figure 5 This is a schematic diagram of the structure of a storage device provided in an embodiment of this application;
[0062] Figure 6This is a schematic diagram illustrating the principle of a large model inference method provided in an embodiment of this application;
[0063] Figure 7 This is a schematic diagram of a computing task that is offloaded to a storage component for execution, provided in an embodiment of this application;
[0064] Figure 8 This is a schematic diagram illustrating an embodiment of the present application for unloading a computing task;
[0065] Figure 9 This is a schematic diagram illustrating a method for implementing large-model inference using a heterogeneous inference system, as provided in an embodiment of this application.
[0066] Figure 10 This is a schematic diagram of the functional architecture of a heterogeneous inference system provided in an embodiment of this application;
[0067] Figure 11 This is a flowchart of a large model inference method provided in an embodiment of this application;
[0068] Figure 12 This is a schematic diagram of the structure of a computing device for performing large model inference, provided in an embodiment of this application;
[0069] Figure 13 This is a schematic diagram of the structure of a storage device for performing large model inference, provided in an embodiment of this application. Detailed Implementation
[0070] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the inference requests and inference data involved in this application are obtained under fully authorized conditions.
[0071] To facilitate understanding, the key terms and concepts involved in this application will be explained below.
[0072] Artificial intelligence (AI) models are a class of mathematical algorithm models that use machine learning concepts to solve practical problems. Typically, AI models include a large number of parameters and calculation formulas (or calculation rules).
[0073] Large models refer to AI models with a large number of parameters and complex computational structures. These models are typically built from deep neural networks and have billions or even hundreds of billions of parameters. Large models have wide applications in various fields, including natural language processing, computer vision, speech recognition, and recommendation systems, but are not limited to these.
[0074] Generative large models are used to perform inference based on a given inference request and generate an inference result. For example, an inference request might instruct inference based on text, images, audio, or video to generate an inference result. Generative large models include large language models (LLMs), multimodal models, and so on.
[0075] Taking generative large-scale models as an example of AI models based on the Transformer architecture, these models typically include multiple Transformer layers, also known as Transformer blocks. A Transformer block refers to the basic building block in a Transformer model; multiple Transformer blocks are stacked sequentially to form a complete Transformer model, enabling effective feature extraction and representation learning from the input sequence. For example, refer to... Figure 1 , Figure 1 This is a schematic diagram of the Transformer layer structure in a large model. For example... Figure 1 As shown, the Transformer layer includes components such as multi-head self-attention, residual connection, layer normalization, and feed-forward network (FFN).
[0076] Multi-head self-attention is a core component of the Transformer architecture. It enables the model to learn dependencies between input sequences (such as words, letters, image features, etc.) in parallel from different representation subspaces, thereby enhancing the model's expressive power. This is also known as the attention computation task of the Transformer layer. Figure 1 As shown, the structure of the multi-head self-attention mechanism mainly includes the following parts:
[0077] Query-key-value (QKV) matrix generation refers to transforming an input sequence into query (Q), key (K), and value (V) matrices through three linear transformations. Each of the three linear transformations corresponds to a different weight matrix, denoted as W. Q WK and W V .
[0078] A multi-head attention layer (or simply attention layer) refers to dividing the Q, K, and V matrices into h attention heads for parallel computation (h is a positive integer). Each attention head focuses on different aspects of the input sequence, thereby capturing richer information. This part of the computation is characterized by bandwidth intensity. For any attention head, its attention score (Score(Q×K)) is calculated. T (That is, multiplying the Q matrix by the transpose of the K matrix), normalize the attention score (achieved through the Softmax normalization operation, denoted as Softmax(S)), multiply the normalization result by the V matrix to obtain the attention output of a single attention head, also known as the context vector Cotext(S×V).
[0079] Linear projection refers to concatenating the attention outputs of multiple attention heads and then performing a linear projection on the concatenated matrix to obtain the output result of the multi-head self-attention mechanism. The weight matrix corresponding to linear projection can be represented as W. O .
[0080] After the output is calculated using the multi-head self-attention mechanism, it is sequentially processed through residual connections and layer normalization before being input into the FFN. The FFN consists of two linear layers (e.g., FF1 and FF2) and a non-linear activation function (usually ReLU). The FFN is used to further transform features and extract information from the layer-normalized output, enhancing the model's expressive power. Residual connections are then added to the output of the FFN, which is then summed with the previously layer-normalized input before undergoing another layer normalization operation to obtain the output of the Transformer layer.
[0081] Furthermore, the inference process of large models typically includes two inference stages: a prefill stage and a decoding stage. The Transformer layers in both the prefill and decoding stages include those described above. Figure 1The components are shown. The prefilling stage is also called the full inference stage, and the decoding stage is also called the incremental inference stage. The prefilling stage is used to process the input sequence of the large model to obtain a sequence of context vectors corresponding to the input sequence. Each vector corresponds to the global semantic representation of a token in the input sequence. The decoding stage is used to generate tokens for the output sequence, i.e., output the inference result, based on the context vector sequence corresponding to the input sequence. The decoding stage is an iterative process, generating a new token in each iteration. For example, if the inference request is the sentence "How is the weather today?", the large model outputs three vectors after the prefilling stage, each vector representing the contextual meaning of "today", "weather", and "how is it" in the entire sentence. These vectors are used in the decoding stage to progressively generate the answer "The weather is sunny today". Typically, the time interval between generating two adjacent tokens in the decoding stage is denoted as "Time Between Tokens" (TBT), used to measure the inference speed of the large model. The smaller the TBT, the faster the model outputs the inference result, and the higher the inference efficiency.
[0082] An accelerator, also known as an acceleration chip, acceleration device, acceleration card, or computing card, is a specialized hardware device or computer system designed to accelerate computation in AI scenarios. In this application, accelerators may include, for example, graphics processing units (GPUs), neural network processing units (NPUs), intelligent processing units (IPUs), tensor processing units (TPUs), or domain-specific architecture (DSA) chips, and are not limited to these.
[0083] An operator (OP) is a computational unit or function that runs on a computing device. In the field of deep learning, neural network layers and even the entire model are composed of operators, which correspond to the computational logic within the neural network layers. For example, a convolutional layer is an operator; the weight summation process in a fully-connected layer (FClayer) is also an operator.
[0084] Processing-in-memory (PIM) is a technique that performs computations inside or near a memory chip. Illustratively, PIM technology involves the following implementations:
[0085] The first type, near data processing (NDP), refers to the technology of integrating computing units on the periphery of memory chips or in adjacent physical layers (such as logic layers in 3D stacking). In NDP technology, computing units and storage units are separate, but physically close to each other to reduce data transmission latency. Storage units primarily provide data access functions, while computing units are located near storage units to perform data processing and computation. This design reduces data transfer between storage and computing units, improving computational efficiency, but the storage units themselves do not participate in computation. Computing units can be, for example, general-purpose processor cores, reconfigurable logic arrays, microprocessor units (MPUs), arithmetic logic units (ALUs), control units (CUs), etc., and are not limited to these. Illustratively, high-bandwidth technologies such as compute express link (CXL), high-bandwidth memory (HBM), and 3D stacking are used to integrate computing units into storage components, giving the storage components computational capabilities.
[0086] For example, in HBM chips employing PIM technology, multiple dynamic random access memory (DRAM) chips are vertically stacked using through silicon via (TSV) technology. Each DRAM chip integrates a computing unit (e.g., logic circuitry deployed around the memory array). These computing units interact with the memory banks via an on-chip bus, thus enabling the HBM chip to perform computations. A bank is a logical grouping of the memory array. Dividing the memory array into multiple banks allows parallel access to data in different banks, thereby improving the bandwidth and efficiency of the DRAM chip. For example, a 16GB DRAM chip includes eight banks, each managing a 2GB memory array area. In some scenarios, each DRAM chip contains multiple bank groups (BGs), each containing multiple banks. Banks within each group share some control circuitry (e.g., row address decoders). Different bank groups operate independently, achieving continuous burst data transfer through inter-bank pipeline operations.
[0087] The second type, in-memory computing, refers to computing using the physical characteristics (such as resistance and charge) of storage units (e.g., newer non-volatile memory). In other words, the storage unit itself possesses computing capabilities and can participate in data processing and computation. For example, resistive random access memory (ReRAM) can simulate vector-matrix operations through changes in resistance.
[0088] The application scenarios of this application are described below.
[0089] This application applies to scenarios where large-model inference is performed using a heterogeneous inference system. The heterogeneous inference system includes a processor and a storage component with computing capabilities. The processor and storage component jointly execute inference requests for large models, instructing inference based on text, images, audio, or video to generate inference results. Indicatively, the processor includes at least one of a central processing unit (CPU) and an accelerator, such as a GPU, NPU, IPU, TPU, DSA chip, etc. In scenarios where the heterogeneous inference system includes multiple types of processors, these processors can be collectively referred to as XPUs. For example, a heterogeneous inference system includes a CPU, an NPU, and a storage component with computing capabilities, where the CPU acts as the main controller, responsible for the allocation and management of overall tasks. The NPU handles AI-related computationally intensive tasks and offloads some computational tasks to the storage component for execution, reducing long-distance data transfer between the NPU and the storage component.
[0090] In related technologies, taking generative large models using the Transformer architecture as an example, when using a heterogeneous inference system to execute inference requests for large models, the computational tasks of the attention layer in the Transformer layer can be offloaded to storage components with computational capabilities. Thus, during the decoding phase of large model inference, when the storage component performs the computational tasks of the attention layer to generate new tokens, it can quickly obtain the KV matrix of the generated tokens, reducing long-distance data transmission. However, storing the KV matrix requires a large amount of storage space in the storage component. When the input sequence of the large model is long or when batch processing the input sequence of the large model, the storage component cannot provide sufficient storage space, thereby limiting the throughput of large model inference and resulting in low inference efficiency. For example, in HBM chips using PIM technology, the bank capacity is limited, and the memory capacity required to store the KV matrix is large and increases proportionally with the batch size.
[0091] Based on this, this application provides a large-model inference method applied to heterogeneous inference systems. For at least one inference request of a large model, the processor determines the execution strategy for at least one inference request to indicate the computational tasks that need to be offloaded to the storage component for execution. The processor and the storage component with computing capabilities then jointly execute the at least one inference request according to the execution strategy. During this process, the processor can sense the memory load of the storage component and adjust the execution strategy in a timely manner based on the memory load of the storage component. For example, a computational task that originally needed to be offloaded to the storage component can be scheduled to be executed on the processor side. In this way, the storage component can release the storage space occupied by the relevant data of the computational task in a timely manner. In this way, the execution strategy can be adjusted in a timely manner according to the memory load of the storage component during the large-model inference process, which can rationally allocate computing resources, effectively improve the resource utilization of the system, adapt to different inference needs, and enable the large model to quickly output inference results when processing long sequences or batch processing input sequences, thereby improving the throughput of large-model inference, that is, improving the efficiency of large-model inference.
[0092] The implementation environment of this application is described below.
[0093] Figure 2 This is a schematic diagram of an implementation environment provided in an embodiment of this application. For example... Figure 2 As shown, the implementation environment includes a heterogeneous inference system 200, which includes a processor 201 and a storage component 202. The processor 201 and the storage component 202 are connected via a bus. For example, the processor 201 and the storage component 202 can be connected via a peripheral component interconnect express (PCIe) link, a Huawei cache coherent system (HCCS) bus, CXL, or NVIDIA Link, etc., and this application does not limit the specific connection.
[0094] The number of processors 201 can be one or more. This application does not limit the type of processor 201. In some embodiments, processor 201 includes at least one of CPU and accelerator. Accelerators are, for example, GPUs, NPUs, IPUs, TPUs, DSA chips, etc. In some embodiments, accelerators are also referred to as accelerator cards, accelerator devices, accelerator chips, computing cards, inference cards, etc., and this application is not limited thereto. In scenarios where the heterogeneous inference system 200 includes multiple types of processors, these processors can be collectively referred to as XPUs. Among them, the CPU acts as the main controller, responsible for the allocation and management of overall tasks. The accelerator handles AI-related computationally intensive tasks and offloads some computational tasks to the storage component for execution, reducing long-distance data transfer between the accelerator and the storage component.
[0095] Storage component 202 includes storage units and computing units, and the storage units are integrated with the computing units, enabling storage component 202 to have computing capabilities. In the embodiments of this application, storage component 202 includes one or more storage units, and each storage unit is integrated with one or more computing units. Multiple storage units can be divided into multiple storage unit groups, and each storage unit group includes one or more storage units. This application does not limit the number of storage units in storage component 202, the number of computing units integrated in each storage unit, the number of storage unit groups, or the number of storage units in a storage unit group. In practical applications, it can be configured according to business needs.
[0096] In storage component 202, storage units provide data storage functions, such as storage banks, meaning memory is partitioned into banks. Computing units provide computing functions, such as general-purpose processor cores, reconfigurable logic arrays, MPUs, ALUs, CUs, etc., and are not limited to these. Integration of storage units and computing units means that computing units are integrated within storage units; correspondingly, storage component 202 is a storage product with in-memory computing capabilities. Alternatively, integration of storage units and computing units means that computing units are integrated at the periphery of storage units or in adjacent physical layers; correspondingly, storage component 202 is a storage product with near-data processing capabilities. For a detailed explanation of in-memory computing and near-data processing, please refer to the foregoing content; it will not be repeated here.
[0097] In some embodiments, the storage component 202 is a memory module, a storage device, or a storage chip integrated on the motherboard of an electronic device, etc., and this application does not limit this to any particular type. In the heterogeneous inference system 200, the number of storage components 202 can be one or more. Figure 2 This is illustrated using multiple storage components 202 as an example. For instance, multiple storage components 202 constitute a storage device.
[0098] In this embodiment, the heterogeneous inference system 200 can access a wired or wireless network to provide large model inference services. The processor 201 and storage component 202 jointly execute inference requests for large models. As described above, the inference process of a large model typically includes a pre-filling stage and a decoding stage. The pre-filling stage is used to process the input sequence of the large model (obtained based on the inference request) to obtain the context vector of each token in the input sequence, providing initial context information for the decoding stage. The decoding stage then generates the tokens of the output sequence, i.e., the output inference result, based on the context vector of each token in the input sequence.
[0099] Based on this, in the heterogeneous inference system 200 provided in this application, the processor 201 determines the execution strategy for at least one inference request of a large model to indicate which computational tasks (such as the computational process of the Transformer layer involved in generating each token) need to be offloaded to the storage component 201 during the execution of at least one inference request. The processor 201 and the storage component 202 then jointly execute the at least one inference request according to the execution strategy. During this process, the processor 201 can sense the memory load of the storage component 202 and adjust the execution strategy accordingly. For example, a computational task that originally needed to be offloaded to the storage component 202 can be scheduled to be executed on the processor 201, thus allowing the storage component 202 to promptly release the storage space occupied by the data related to that computational task. In this way, the execution strategy can be adjusted in a timely manner according to the memory load of the storage component 202 during the large model inference process. This can reasonably allocate computing resources, effectively improve the resource utilization of the system, adapt to different inference needs, and enable the large model to quickly output inference results when processing long sequences or batch processing input sequences, thereby improving the throughput of large model inference, that is, improving the efficiency of large model inference.
[0100] Furthermore, the heterogeneous inference system 200 can be deployed on general-purpose physical servers, desktop computers, terminal devices, etc. For example, see reference... Figure 3 , Figure 3 This is a schematic diagram of a heterogeneous inference system provided in an embodiment of this application. For example... Figure 3 As shown, taking the heterogeneous inference system 200 deployed in a server as an example, the processor 201 includes the CPU and accelerators (such as NPUs) of the host in the server, and the storage component 202 is HBM memory. The CPU, NPU, and HBM memory are connected and interact with each other via a bus (such as a PCIe link). It should be understood that a host is a complete device or module including a CPU, memory, and other related components, which works with the NPU and HBM memory in the server to achieve various complex computing functions. When there are multiple NPUs in the server, the NPUs communicate with each other through high-speed interconnect links, such as HCCS, RDMA over converged Ethernet (RoCE), NvLink, CXL, universal chiplet interconnect express (UCIe), cache coherent interconnect for accelerators (CCIX), etc., and this application is not limited to these.
[0101] In some embodiments, the heterogeneous inference system 200 is deployed on a cloud platform. A cloud platform, short for cloud computing platform, refers to a service that provides computing, networking, and storage capabilities based on hardware and software resources. Through the network "cloud," massive amounts of data are processed and analyzed remotely before being returned to the user. It features large scale, distributed architecture, virtualization, high availability, scalability, on-demand service, and security. Cloud platforms can achieve rapid provisioning and release of configurable computing resources with relatively low management costs or low interaction complexity between users and service providers.
[0102] The aforementioned wireless or wired networks utilize standard communication technologies and / or protocols. These networks are typically Transmission Control Protocol / Internet Protocol (TCP / IP) networks used in data center networks, as well as RDMA networks such as RoCE networks and InfiniBand (IB) networks; no limitation is made thereto. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0103] Based on the above Figure 2 and Figure 3 In the implementation environment shown, this application provides a computing device capable of implementing the functions of the processor 201 in the heterogeneous inference system 200 described above. (Refer to the following...) Figure 4 The structure of the computing device will be introduced. Figure 4 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Figure 4 As shown, the computing device 400 includes a memory 401, a processor 402, a communication interface 403, and a bus 404. The memory 401, processor 402, and communication interface 403 are interconnected via the bus 404.
[0104] The memory 401 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or it may be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. In this application, the memory 401 is used to store at least one piece of program code. When the program code stored in the memory 401 is executed by the processor 402, the processor 402 is used to execute the large model inference method provided in this application.
[0105] The processor 402 may be a network processor (NP), a CPU, an application-specific integrated circuit (ASIC), or an integrated circuit used to control the execution of the program in this application. The processor 402 may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The number of processors 402 may be one or more. For example, depending on the actual needs of large-scale model inference, the computing device 400 may include multiple processors 402, which may include at least two of the following: CPU, GPU, NPU, IPU, TPU, DSA chip, etc. This application does not limit the number of processors in this regard.
[0106] The communication interface 403 uses a transceiver module, such as a transceiver, to enable communication between the computing device 400 and other devices or communication networks. For example, data can be acquired through the communication interface 403. It can also be connected to a storage component.
[0107] The memory 401 and the processor 402 can be set separately or integrated together.
[0108] Bus 404 may include a pathway for transmitting information between various components of computing device 400 (e.g., memory 401, processor 402, communication interface 403).
[0109] This application also provides a storage device capable of implementing the functions of the storage component 202 in the heterogeneous inference system 200 described above. (Refer to the following...) Figure 5 The structure of storage devices will be introduced. Figure 5 This is a schematic diagram of the structure of a storage device provided in an embodiment of this application. Figure 5 As shown, the storage device 500 includes a storage unit 501, a computing unit 502, a communication interface 503, and a bus 504. The storage unit 501, the computing unit 502, and the communication interface 503 are interconnected via the bus 504.
[0110] There are multiple storage units 501. The diagram illustrates an example where a computing unit 502 is integrated around a storage unit 501. The storage unit 501 provides data storage functionality, such as a storage bank.
[0111] The computing unit 502 is used to provide computing functions, such as a general-purpose processor core, a reconfigurable logic array, an MPU, an ALU, a CU, etc., and is not limited thereto. In some embodiments, the computing unit 502 includes a vector unit and a matrix-vector multiplication unit, wherein the vector unit is used to provide computing power for operations between vectors, i.e., vector computing power; the matrix-vector multiplication unit is used to provide computing power for multiplication between matrices and vectors, i.e., general matrix-vector multiplication (GEMV) computing power.
[0112] Communication interface 503 is used to enable communication between storage device 500 and other devices or communication networks using a transceiver module such as a transceiver. For example, data can be acquired through communication interface 503. Alternatively, it can be connected to a processor via communication interface 403.
[0113] Bus 504 may include a path for transmitting information between various components of storage device 500 (e.g., storage unit 501, computing unit 502, communication interface 503).
[0114] It should be noted that the above Figure 4 and Figure 5 The figures shown are structural diagrams of a computing device and a storage device provided in this application. In some embodiments, the computing device 400 may include other components to achieve more functions. Similarly, the storage device 500 may also include other components to achieve more functions. This application does not limit this. In other scenarios, the above-mentioned computing device and storage device may also be integrated together, for example, both deployed in a physical server. This application does not limit this.
[0115] The large-scale model inference method provided in this application is described below.
[0116] For ease of understanding, please refer to the following: Figures 6 to 8 This paper introduces the principles of large-scale model reasoning methods.
[0117] Figure 6 This is a schematic diagram illustrating the principle of a large model inference method provided in an embodiment of this application. For example... Figure 6 As shown, this method is applied to a heterogeneous inference system, which includes processors and storage components. The large model inference method involves the following stages: system initialization, load-aware policy determination, and execution of inference requests according to the policy. These stages are described in detail below.
[0118] System initialization phase.
[0119] During the system initialization phase, the processor determines an initial execution strategy based on at least one inference request from the large model. This execution strategy indicates the computational tasks that need to be offloaded to the storage component during the execution of the at least one inference request. In some embodiments, the execution strategy also indicates the computational tasks that need to be scheduled to the processor during the execution of the at least one inference request. In other words, the execution strategy can indicate the computational tasks that the processor and the storage component each need to perform during the execution of the at least one inference request.
[0120] In this application, the computational tasks offloaded to the storage component are used to perform the computation process of the Transformer layer in the large model. The structure of the Transformer layer in the large model is described above. Figure 1 The details shown are not repeated here. Illustratively, in this application, the computational tasks offloaded to the storage component include at least one of the attention computation task and the feedforward network computation task of the Transformer layer in the large model; wherein, the attention computation task is used to execute a multi-head self-attention mechanism, including at least one of the following: query-key-value QKV matrix generation task, attention score computation task, normalization computation task, context vector computation task, and linear projection task. It should be understood that the Transformer layer is a core component of the large model. Offloading the Transformer layer's computational tasks to the storage component allows data to be processed internally, significantly reducing data transfer time and thus improving task execution efficiency. For example, in the attention score computation task, computation within the storage component avoids transferring a large number of intermediate results to the processor, thereby alleviating memory bandwidth pressure. Moreover, offloading both the attention computation task and the feedforward network computation task to the storage component enables full-process offloading of the Transformer layer, maximizing the reduction of data transfer time and improving the inference efficiency of the large model.
[0121] Furthermore, in practical applications, the computational tasks to be offloaded to storage components can be flexibly configured based on the processor and storage component ratio in a heterogeneous inference system. For example, the processor determines the initial execution strategy based on the number of at least one inference request, the number of storage components, the number of processors, the number of storage units in the storage components, and the number of attention heads in the Transformer layer of the large model. Here, one storage unit in a storage component corresponds to at least one attention head in the Transformer layer of the large model. In this way, the processor can comprehensively determine the initial execution strategy based on multiple factors, and the correlation between storage units and the number of attention heads in the Transformer layer allows for the rational allocation of the correspondence between storage units in storage components and attention heads in the Transformer layer, ensuring the reasonable allocation and effective utilization of computing resources and improving the overall performance and resource utilization of large model inference.
[0122] For example, refer to Figure 7 , Figure 7 This is a schematic diagram illustrating a computing task that is offloaded to a storage component for execution, as provided in an embodiment of this application. Figure 7 As shown, this example illustrates a heterogeneous inference system that batch-processes multiple inference requests for a large model, with the storage component comprising multiple storage unit groups. Each storage unit group includes multiple storage units, which are integrated with the computation unit. The heterogeneous inference system can achieve multi-dimensional parallel Attention-FFN full offloading. Attention refers to the attention computation task of the Transformer layer, and FFN refers to the feedforward network computation task of the Transformer layer. Illustratively, multi-dimensional parallel Attention-FFN full offloading involves the following points:
[0123] (1) Divide multiple inference requests into multiple sub-batches and use multiple memory groups for parallel computation. For example, divide 10 inference requests into 2 sub-batches. Sub-batches 0 are computed using memory group 0 (as shown in bank group 0 in the figure), and sub-batches 1 are computed using memory group 1 (as shown in bank group 1 in the figure).
[0124] (2) For any group of storage units, the attention computation task is offloaded to a single storage unit for execution according to the number of attention heads in the Transformer layer. For example, if a group of storage units includes 4 storage units and the number of attention heads is 32, then each storage unit corresponds to 8 attention heads in the Transformer layer. Accordingly, taking sub-batch 1, which includes 5 inference requests, as an example, one storage unit corresponds to 8 attention heads for each of the 5 inference requests. That is, each storage unit independently processes its allocated attention heads, and the attention computation is performed in parallel using on-chip computing units. The attention heads of different inference requests can be pipelined or processed concurrently to improve computational efficiency.
[0125] (3) For any attention head corresponding to any storage unit, the attention calculation tasks corresponding to that attention head include query-key-value QKV matrix generation, attention score calculation, normalization calculation, context vector calculation, and linear projection. The model weights for the QKV matrix generation task include the weight matrices corresponding to the QKV matrices, which are then split by columns and stored in the storage unit. The weight matrices corresponding to the linear projection task are split by rows and stored in the storage unit.
[0126] (4) For any inference request, the attention outputs of the multiple attention heads corresponding to the inference request are concatenated (or aggregated) to obtain the execution result of the attention computation task for that inference request. Schematic, this process is performed by the XPU, for example, by an accelerator.
[0127] (5) For the feedforward network computation task, the model weights of the two linear layers in the FFN are partitioned by column and by row, respectively, and then stored in each storage unit. Based on this, for any inference request, the feedforward network computation task can be executed by each storage unit in the storage unit group according to the execution result of the attention computation task. In some scenarios, a portion of the model weights of the two linear layers in the FFN are partitioned by column and by row, respectively, and then stored in each storage unit, while the remaining model weights are offloaded to the XPU. That is, the feedforward network computation task is jointly executed by the XPU and the storage component. In other scenarios, the model weights of the two linear layers are stored in the XPU, and the XPU executes the feedforward network computation task. This application does not limit this.
[0128] As can be seen, through the above method, the heterogeneous inference system can achieve multi-dimensional parallel Attention-FFN full offloading, thereby making full use of the computing resources of the storage component. The system initialization phase involves orchestrating pipeline strategies based on model information (such as the number of Transformer layers and the number of attention heads in each Transformer layer) and hardware information (such as the computing power of the XPU and storage components, memory bandwidth, cache size, etc.). This is achieved by creating kernel programs or subgraph partitioning to generate execution strategies for multiple inference requests. After constructing and compiling the offloading kernel functions according to the device drivers and corresponding instruction sets in the heterogeneous inference system, the computational tasks are scheduled to be executed on the XPU and / or storage components.
[0129] Load-aware strategy determination phase.
[0130] After the heterogeneous inference system executes the inference request according to the initial execution strategy determined by the processor, it enters the load-aware strategy determination stage. In this stage, the processor's memory management unit can sense the memory load of the storage components and adjust the execution strategy in a timely manner according to the memory load of the storage components. For example, it can schedule a computing task that originally needed to be offloaded to the storage components to be executed on the processor side.
[0131] For example, during the decoding phase of an inference request, as new tokens are generated, the available memory of the storage component gradually decreases. Based on this, the processor can determine weight reclamation when generating each token, based on the memory changes in the storage component. This means determining whether the model weights stored on the storage component need to be reclaimed to the processor and adjusting the execution strategy accordingly. Specifically, the processor first determines whether the model weights corresponding to the feedforward network computation task need to be reclaimed (e.g., whether the memory load of the storage component meets the requirements after reclaiming 80% of the model weights). If the memory load of the storage component still does not meet the requirements after reclaiming all the model weights corresponding to the feedforward network computation task, then it determines whether the model weights corresponding to at least one of the QKV matrix generation task and the linear projection task need to be reclaimed. Based on the determined model weights that need to be reclaimed, a memory reclamation strategy is generated, and the model weights stored on the storage component are unloaded according to the memory reclamation strategy. It should be understood that since the feedforward network computation task is implemented through matrix multiplication operators and is a computationally intensive task, determining whether the model weights corresponding to the feedforward network computation task need to be reclaimed first allows computationally intensive tasks to be prioritized for execution on the XPU side, improving computational efficiency. It should be understood that this introduction uses model weights as an example. In some embodiments, model weights and input data of the computation task can also be judged together, with the same principle, so it will not be described again.
[0132] The phase of executing the inference request according to the strategy.
[0133] After the processor determines the execution strategy for the inference request, the processor and storage component jointly execute the inference request according to the execution strategy. Specifically, the computation units in the storage component execute computational tasks offloaded to the storage component based on the inference data stored in the storage unit. The inference data includes at least one of the model weights and the input data for the computational tasks. Taking a processor consisting of a CPU and an accelerator as an example, the CPU controls the accelerator to execute tasks in the inference request other than the computational tasks that need to be offloaded to the storage component, according to the execution strategy. Furthermore, the CPU controls the accelerator to offload computational tasks that need to be offloaded to the storage component to the storage component for execution, according to the execution strategy. For example, computational tasks offloaded to the storage component include attention computation tasks and feedforward network computation tasks in the Transformer layer of a large model; where attention computation tasks include QKV matrix generation, attention score calculation, normalization calculation, context vector calculation, and linear projection. In the figure, H0 and H1 represent different attention heads. Additionally, the figure illustrates an example where the storage component and processor jointly execute FFN computation tasks, and the processor concatenates the attention outputs of multiple attention heads.
[0134] Furthermore, after determining the memory reclamation strategy and the adjusted execution strategy through the load-aware strategy determination phase, on the one hand, the processor sends a memory reclamation request to the storage component according to the memory reclamation strategy, and the storage component unloads the inference data on the storage component according to the instructions of the memory reclamation request; on the other hand, the processor and the storage component continue to jointly execute the inference request according to the adjusted execution strategy. It can be seen that in the large model inference method provided in this application, based on memory load and memory awareness, and while ensuring sufficient cache space in the storage component, dynamically shifting some model weights and scheduling attention calculation tasks and feedforward network calculation tasks can fully utilize the idle computing power of the storage component. For example, the storage component and the XPU jointly execute feedforward network calculation tasks, thereby effectively improving model inference efficiency.
[0135] Indicatively, for reference Figure 8 , Figure 8 This is a schematic diagram of an embodiment of the present application for unloading a computing task.
[0136] like Figure 8As shown in Figure (a), in the heterogeneous inference system provided in this application, for any given storage unit, the storage unit corresponds to at least one attention head. Taking one attention head as an example (H0 represents a single attention head), the computation unit integrated in the storage unit can execute the attention computation task corresponding to that attention head. Specifically, the vector unit in the computation unit is used to perform the normalization computation task, and the matrix-vector multiplication unit in the computation unit is used to perform the QKV matrix generation task, the attention score calculation task, the context vector calculation task, and the linear projection task. Furthermore, the storage unit can also store some model weights for the feedforward network computation task; correspondingly, the matrix-vector multiplication unit in the computation unit executes part of the feedforward network computation task.
[0137] like Figure 8 As shown in Figure (b), in the heterogeneous inference system provided by the relevant technology, for the computation process of the Transformer layer in a large model, each operator is usually executed serially on the XPU. When it is necessary to unload the computation task to the storage component for execution, only the attention score calculation task, normalization calculation task, and context vector calculation task in the attention calculation task are unloaded to the storage component for execution, resulting in a large amount of idle computing power in the storage component.
[0138] Further, refer to Figure 9 , Figure 9 This is a schematic diagram illustrating a method for implementing large-model inference using a heterogeneous inference system, as provided in an embodiment of this application. Figure 9 As shown, taking a single storage unit bank in the storage component as an example, this storage unit stores the KV matrix corresponding to a single attention head, the model weights corresponding to at least one of the QKV matrix generation task and the linear projection task, and a portion of the model weights corresponding to the feedforward network computation task (or all model weights; here, a portion of the model weights is used as an example). The processor (i.e., the XPU side) can sense the memory load of the storage component, adjust the execution strategy and generate a corresponding memory reclamation strategy based on the memory load, and then move and schedule tasks according to the weights of the memory reclamation strategy, and schedule related tasks according to the adjusted execution strategy.
[0139] As can be seen, compared with related technologies, the large model inference method provided in this application can make full use of the idle computing power of the storage component, thereby improving the model inference efficiency.
[0140] The following is for reference. Figure 10 In conjunction with the heterogeneous reasoning system described in the aforementioned implementation environment, the principle of the large-model reasoning method provided in this application will be further introduced. Figure 10 This is a schematic diagram of the functional architecture of a heterogeneous inference system provided in an embodiment of this application. For example... Figure 10As shown, the heterogeneous reasoning system is used to provide a strategy determination function 1001, a strategy execution function 1002, and a strategy adjustment function 1003.
[0141] The strategy determination function 1001 is implemented by the processor in the heterogeneous inference system, specifically by the XPU. Illustratively, the processor determines an execution strategy based on at least one inference request from a large model. This execution strategy indicates the computational tasks that need to be offloaded to the storage component during the execution of the at least one inference request. These computational tasks are used to perform computations in the Transformer layer of the large model. Each computational task offloaded to the storage component can be implemented by a function or an operator. Furthermore, when determining the execution strategy, the processor can flexibly set the granularity of the tasks related to the inference request across multiple dimensions, based on model and hardware information, while meeting inference latency requirements.
[0142] The policy execution function 1002 is jointly implemented by the processor and storage components in the heterogeneous inference system. As described above, the execution policy determined by the processor based on at least one inference request can instruct the computational tasks that the processor and storage components must each perform during the execution of at least one inference request. Based on this, the processor and storage components jointly execute at least one inference request according to the execution policy. Specifically, the computational unit of the storage component executes computational tasks offloaded to the storage component based on inference data stored in the storage unit. The inference data includes at least one of model weights and input data for the computational tasks. In some embodiments, the processor includes a CPU and an accelerator. The CPU controls the accelerator to execute tasks in the inference request other than the computational tasks to be offloaded to the storage component, according to the execution policy. Furthermore, the CPU controls the accelerator to offload computational tasks to the storage component for execution, according to the execution policy.
[0143] The strategy adjustment function 1003 is implemented by the processor in the heterogeneous inference system. Illustratively, during the execution of at least one inference request by the processor and storage components according to the execution strategy, the processor adjusts the execution strategy in a timely manner based on the execution strategy and the memory usage status of the storage components. The memory usage status of the storage components indicates their memory load. For example, during the decoding phase of at least one inference request, as new tokens are generated, the available memory of the storage components gradually decreases. Based on this, the processor can, when generating each token, determine whether to reclaim model weights stored on the storage components to the processor side, based on the memory usage of the previous token generation process, according to the memory usage status of the storage components. If it is determined that model weights need to be reclaimed, a corresponding memory reclamation strategy is generated and the execution strategy is adjusted accordingly. After generating the memory reclamation strategy and the adjusted execution strategy, the strategy execution function 1002 reclaims the inference data on the storage components according to the memory reclamation strategy, and continues to jointly execute at least one inference request according to the adjusted execution strategy.
[0144] Furthermore, the functional division of heterogeneous reasoning systems is not limited to... Figure 10 As shown, for example, the heterogeneous inference system is also used to provide storage functions, which are jointly implemented by the processor and storage components in the heterogeneous inference system. In practical applications, more functions can be set according to user needs, which is not limited in this application.
[0145] Schematic illustration: the functionality provided by the aforementioned heterogeneous inference system can be installed as a software toolkit component on a server and run by the server's processor and storage components. For example, when the server's hardware architecture employs a heterogeneous computing architecture (including computing architectures using computing units with different types of instruction sets), users can install heterogeneous computing frameworks on the server. One such framework is the Compute Architecture for NeuroNet (CANN), a heterogeneous computing framework for neural networks. CANN can support users in quickly building AI applications by providing multi-layered programming interfaces. Additionally, users can install deep learning frameworks on the server to compile methods for implementing models, construct large-scale computational graphs, and automatically perform gradient calculations within the computational graphs. The functionality provided by the aforementioned heterogeneous inference system can be interfaced and adapted with deep learning frameworks and heterogeneous computing frameworks.
[0146] The following is for reference. Figure 11 The illustrated method embodiment describes the flow of the large model inference method provided in this application.
[0147] Figure 11This is a flowchart of a large model inference method provided in an embodiment of this application. For example... Figure 11 As shown, this method is applied to a heterogeneous inference system, which includes a processor and a storage component. The storage component includes storage units and computation units, and the storage units and computation units are integrated. This large model inference method includes the following steps 1101 to 1104.
[0148] 1101. The processor determines a first execution strategy based on at least one inference request from the large model.
[0149] In this application's embodiments, a large model refers to a generative large model employing the Transformer architecture. The inference request for the large model includes user input information, which can be text, voice, images, video, etc., and this application does not limit this. In some embodiments, the inference request also includes other information, such as the version of the large model, user information, request type (e.g., image recognition or text processing), etc., and is not limited thereto. This application does not limit the number of inference requests processed in parallel by the heterogeneous inference system.
[0150] In this embodiment, the initial execution strategy determined by the processor based on at least one inference request of the large model is referred to as the first execution strategy, and the computational tasks that need to be offloaded to the storage component as indicated by the first execution strategy are referred to as the first partial computational tasks (the first execution strategy and the first partial computational tasks can also use other naming conventions, which are not limited in this application). Illustratively, the first execution strategy indicates the first partial computational tasks that need to be offloaded to the storage component during the execution of at least one inference request, wherein the computational tasks offloaded to the storage component are used to perform the computation process of the Transformer layer in the large model. Illustratively, the first execution strategy is determined by the CPU in the processor based on at least one inference request of the large model.
[0151] In some embodiments, the processor determines a first execution strategy based on at least one inference request of the large model, including: determining the first execution strategy based on the number of at least one inference request, the number of storage components, the number of processors, the number of storage units in the storage components, and the number of attention heads in the Transformer layer of the large model, wherein one storage unit in the storage component corresponds to at least one attention head in the Transformer layer of the large model. This process is similar to the aforementioned... Figure 7The same principle applies to the multidimensional parallel Attention-FFN full offloading described earlier. That is, the processor can flexibly configure the computational tasks to be offloaded to the storage components based on the processor-to-storage component ratio in the heterogeneous inference system, the hardware information of the storage components, and the model information of the large model, etc. The computational tasks offloaded to the storage components include at least one of the attention computation tasks of the Transformer layers and the feedforward network computation tasks in the large model; the attention computation tasks include at least one of the following: QKV matrix generation, attention score calculation, normalization calculation, context vector calculation, and linear projection.
[0152] For example, a large model has 10 inference requests. The first execution strategy is as follows: divide the 10 inference requests into two sub-batches. Sub-batches 0 are computed using storage unit group 0, and sub-batches 1 are computed using storage unit group 1. For any storage unit group, based on the Transformer layer's attention heads (32), allocate 8 attention heads for each of the 5 inference requests to one storage unit. For any storage unit, the computational tasks to be offloaded to that storage unit include: attention computation tasks corresponding to a single attention head and parts of the feedforward network computation tasks. The attention computation tasks include QKV matrix generation, attention score calculation, normalization calculation, context vector calculation, and linear projection. The model weights for the QKV matrix generation task include the weight matrices corresponding to the QKV matrices, which are then split by column and stored in the storage units. The weight matrices corresponding to the linear projection task are split by row and stored in the storage units. For the feedforward network computation task, the model weights of the two linear layers in the FFN are split by column and by row, respectively, and stored in the respective storage units.
[0153] 1102. The processor and storage components jointly execute at least one inference request in accordance with the first execution strategy.
[0154] In this embodiment, the processor and storage component execute tasks scheduled to themselves according to the instructions of a first execution strategy. Illustratively, for the storage component, the computing units within the storage component execute computing tasks offloaded to the storage component based on inference data stored in the storage unit. The inference data includes at least one of model weights and input data for the computing tasks. For the processor, the processor executes computing tasks not offloaded to the storage component according to the instructions of the first execution strategy.
[0155] In some embodiments, the computation units in the storage component include vector units and matrix-vector multiplication units. Accordingly, the computation units in the storage component perform computational tasks offloaded to the storage component based on the inference data stored in the storage unit, including: the vector units performing normalization computation tasks offloaded to the storage component based on the inference data; and the matrix-vector multiplication units performing tasks other than the normalization computation tasks offloaded to the storage component based on the inference data. This clearly defined division of labor allows for optimization based on the characteristics of different computational tasks, improving computational parallelism and efficiency, further enhancing the performance of the storage component in performing computational tasks, and making large model inference more efficient.
[0156] Furthermore, since the various storage units within the storage components, and the individual storage units within those units, can execute in parallel, the GEMV computing power provided by multiple matrix-vector multiplication units (i.e., through batch processing) can achieve the same effect as the general matrix multiplication (GEMM) computing power provided by the accelerator. For example, in the decoding stage of large model inference, batch processing allows QKV matrix generation tasks, linear projection tasks, and feedforward network computation tasks to be transformed from bandwidth-intensive to computationally intensive, thereby effectively improving the efficiency of large model inference.
[0157] In some embodiments, the processor includes a CPU and an accelerator, with the CPU controlling the accelerator. The processor and storage components jointly execute at least one inference request according to a first execution strategy, including: the CPU controlling the accelerator to execute tasks in the at least one inference request, excluding computational tasks that need to be offloaded to the storage component, and controlling the accelerator to offload computational tasks to the storage component for execution, according to the first execution strategy. That is, the CPU acts as the main controller, responsible for the overall task allocation and management. This heterogeneous processor architecture fully leverages the control capabilities of the CPU and the computational advantages of the accelerator, rationally allocating computational tasks and offloading tasks other than those executed by the storage component to the accelerator, thus improving computational efficiency. Simultaneously, the CPU's control over the accelerator better coordinates system resources, ensuring the smooth progress of the inference process.
[0158] In some embodiments, the storage component includes multiple storage unit groups, each storage unit group including multiple storage units. Accordingly, the processor and the storage component jointly execute at least one inference request according to a first execution strategy, including: the processor and each storage unit group in the multiple storage unit groups jointly execute multiple inference requests according to the first execution strategy, wherein one storage unit group corresponds to at least one inference request. This grouped execution method enables parallel processing of multiple inference requests, fully utilizing the parallel computing capabilities of the storage component, improving inference efficiency, and accelerating the processing speed of inference requests, making it particularly suitable for scenarios requiring the processing of a large number of inference requests.
[0159] 1103. During the execution of at least one inference request, the processor determines a second execution strategy based on the first execution strategy and the memory usage state of the storage components.
[0160] In this embodiment, during the execution of at least one inference request, the processor can sense the memory load of the storage component and adjust the execution strategy in a timely manner according to the current first execution strategy and the memory usage status of the storage component. In this embodiment, the adjusted execution strategy is named the second execution strategy, and the computational task that needs to be offloaded to the storage component for execution indicated by the second execution strategy is called the second part of the computational task (the second execution strategy and the second part of the computational task can also use other naming methods, which are not limited in this application). Schematic, the second execution strategy indicates the second part of the computational task that needs to be offloaded to the storage component for execution during the execution of at least one inference request.
[0161] In some embodiments, the processor determines a second execution strategy based on a first execution strategy and the memory usage state of the storage component. This includes: the processor determining a second execution strategy and a memory reclamation strategy based on the first execution strategy and the memory usage state of the storage component. The memory reclamation strategy indicates the inference data that needs to be unloaded from the storage component during the decoding phase of at least one inference request. Based on this, the processor sends a memory reclamation request to the storage component according to the memory reclamation strategy; the storage component unloads the inference data according to the instructions of the memory reclamation request. This process means that the processor generates a memory reclamation strategy while determining the second execution strategy, and then reclaims the inference data on the storage component according to the memory reclamation strategy to release the memory space of the storage component, providing technical support for improving the throughput of large models.
[0162] In some embodiments, the process by which the processor determines the second execution strategy and the memory reclamation strategy includes the following steps:
[0163] Step 1: The processor determines first reference information based on the first execution strategy and memory usage status. The first reference information indicates the memory usage of the storage component when continuing to execute at least one inference request according to the first execution strategy.
[0164] Step 2: The processor determines the second execution strategy and memory reclamation strategy based on the first reference information and the type of the first part of the computational tasks indicated by the first execution strategy. The first part of the computational tasks indicated by the first execution strategy includes attention computation tasks and feedforward network computation tasks in the Transformer layer of the large model.
[0165] The processor's execution of step 2 includes: the processor determining second reference information based on the first reference information and the feedforward network computation task indicated by the first execution strategy; the second reference information indicating the memory usage of the storage component after the inference data of the feedforward network computation task is unloaded from the storage component; and the processor determining a second execution strategy and a memory reclamation strategy based on the second reference information and the first execution strategy.
[0166] If the memory usage indicated by the second reference information meets the conditions, the processor determines a second execution strategy and a memory reclamation strategy based on the second reference information and the first execution strategy. If the memory usage indicated by the second reference information does not meet the conditions, the processor determines third reference information based on the attention computation task indicated by the second reference information and the first execution strategy, and determines a second execution strategy and a memory reclamation strategy based on the third reference information and the first execution strategy. The third reference information indicates the memory usage of the storage component after unloading the inference data of the feedforward network computation task and the attention computation task from the storage component. For example, the condition may be that the memory usage of the storage component is less than a threshold. The threshold can be set according to business needs, for example, 90%, and this application does not limit this setting.
[0167] The process described above involves the processor performing multi-stage inference data reclamation checks when generating each token, based on changes in the storage component's memory. For example, it determines whether model weights stored on the storage component need to be reclaimed to the processor and adjusts the execution strategy accordingly. Specifically, the processor first determines whether the inference data corresponding to the feedforward network computation task needs reclamation. If, after reclamation of all inference data for the feedforward network computation task, the memory load of the storage component still does not meet requirements, it then determines whether the inference data corresponding to at least one of the QKV matrix generation task and the linear projection task needs reclamation. Based on the determined inference data requiring reclamation, a memory reclamation strategy is generated, and the inference data stored on the storage component is unloaded according to the strategy. It should be understood that since the feedforward network computation task is implemented through matrix multiplication operators and is a computationally intensive task, determining whether the inference data corresponding to the feedforward network computation task needs reclamation first allows computationally intensive tasks to be prioritized for execution on the XPU side, improving computational efficiency.
[0168] 1104. The processor and storage components continue to jointly execute at least one inference request in accordance with the second execution strategy.
[0169] In this embodiment, after determining the second execution strategy in step 1103, the processor and storage component continue to jointly execute at least one inference request according to the second execution strategy. The computing units in the storage component execute computing tasks offloaded to the storage component based on the inference data stored in the storage unit. For the processor, the processor executes computing tasks not offloaded to the storage component according to the instructions of the second execution strategy.
[0170] In this way, the heterogeneous inference system implements a memory-aware dynamic memory reclamation mechanism during the execution of at least one inference request. Correspondingly, there are multiple combinations of computing tasks executed by the computing units in the storage component. The switching logic of this combination lies in the processor's awareness of the memory load of the storage component, which can minimize the resource idle rate in scenarios where inference requests change dynamically.
[0171] Schematic representation: The above combination methods involve at least one of the following forms:
[0172] (1) Bandwidth-intensive computation process; namely, attention score calculation task, normalization calculation task, and context vector calculation task.
[0173] (2) Bandwidth-intensive and partially computationally intensive computation processes; namely, QKV matrix generation task, attention score calculation task, normalization calculation task, context vector calculation task, and linear projection task.
[0174] (3) The entire calculation process; namely, the QKV matrix generation task, attention score calculation task, normalization calculation task, context vector calculation task, linear projection task, and feedforward network calculation task.
[0175] As can be seen, the above method can fully utilize the heterogeneous computing power in the heterogeneous inference system, maximize the utilization of computing units, and improve TBT in the decoding stage. Moreover, by offloading some of the weights and KV matrices originally on the XPU side to the storage component, the XPU side can increase throughput and reduce memory overhead.
[0176] In summary, the large model inference method provided in this application, for at least one inference request of a large model, the processor determines the execution strategy for at least one inference request to indicate the computational tasks that need to be offloaded to the storage component for execution, and the processor and the storage component with computing power jointly execute at least one inference request according to the execution strategy. During this process, the processor can sense the memory load of the storage component and adjust the execution strategy in a timely manner based on the memory load of the storage component. For example, it can schedule a computational task that originally needed to be offloaded to the storage component to be executed on the processor side, so that the storage component can release the storage space occupied by the relevant data of the computational task in a timely manner. In this way, by adjusting the execution strategy in a timely manner according to the memory load of the storage component during large model inference, computational resources can be allocated reasonably, effectively improving the resource utilization of the system, adapting to different inference needs, and enabling large models to quickly output inference results even when processing long sequences or batch processing input sequences, thereby improving the throughput of large model inference, that is, improving the efficiency of large model inference.
[0177] Based on the above-described large model inference method, this application also provides a computing device for performing large model inference and a storage device for performing large model inference. The computing device can implement some or all of the functions of the processor in the aforementioned heterogeneous inference system through software, hardware, or a combination of both. The storage device can implement some or all of the functions of the storage component in the aforementioned heterogeneous inference system through software, hardware, or a combination of both. References are given below. Figure 12 and Figure 13 This section introduces the structure of computing and storage devices.
[0178] Figure 12 This is a schematic diagram of the structure of a computing device for performing large model inference, provided in an embodiment of this application. Figure 12 As shown, the computing device includes a first determining module 1201, an execution module 1202, and a second determining module 1203.
[0179] The first determining module 1201 is used to determine a first execution strategy based on at least one inference request of the large model. The first execution strategy indicates a first part of the computational task that needs to be offloaded to the storage component for execution during the execution of at least one inference request.
[0180] Execution module 1202 is configured to execute at least one computational task in an inference request that has not been offloaded to the storage component, in accordance with a first execution strategy.
[0181] The second determining module 1203 is used to determine a second execution strategy based on the first execution strategy and the memory usage state of the storage component during the execution of at least one inference request. The second execution strategy indicates a second part of the computational task that needs to be offloaded to the storage component for execution during the execution of at least one inference request.
[0182] The execution module 1203 is also configured to continue executing at least one computational task in an inference request that has not been offloaded to the storage component, in accordance with the second execution strategy.
[0183] In some embodiments, the computational tasks to be offloaded to the storage component include at least one of the attention computation task of the Transformer layer in the large model and the feedforward network computation task; wherein, the attention computation task includes at least one of the QKV matrix generation task, attention score calculation task, normalization calculation task, context vector calculation task and linear projection task.
[0184] In some embodiments, the first determining module 1201 is configured to: determine a first execution strategy based on the number of at least one inference request, the number of storage components, the number of processors, the number of storage units in the storage components, and the number of attention heads in the Transformer layer of the large model; wherein, one storage unit in the storage component corresponds to at least one attention head in the Transformer layer of the large model.
[0185] In some embodiments, the second determining module 1203 is configured to: determine a second execution strategy and a memory reclamation strategy based on a first execution strategy and the memory usage state of a storage component during the execution of the decoding phase of at least one inference request; the memory reclamation strategy indicates the inference data that needs to be unloaded from the storage component during the execution of the decoding phase of at least one inference request.
[0186] In some embodiments, the computing device further includes a sending module 1204, configured to: send a memory reclamation request to the storage component in accordance with a memory reclamation strategy.
[0187] In some embodiments, the second determining module 1203 is configured to: determine first reference information based on a first execution strategy and memory usage status, wherein the first reference information indicates the memory usage of the storage component when at least one inference request is executed in accordance with the first execution strategy; and determine a second execution strategy and a memory reclamation strategy based on the first reference information and the type of the first part of the computation task.
[0188] In some embodiments, the first part of the computation task includes an attention computation task and a feedforward network computation task in the Transformer layer of the large model; the second determining module 1203 is configured to: determine second reference information based on the first reference information and the feedforward network computation task in the first part of the computation task, the second reference information indicating the memory usage of the storage component after unloading the inference data of the feedforward network computation task in the first part of the computation task from the storage component; and determine a second execution strategy and a memory reclamation strategy based on the second reference information and the first execution strategy.
[0189] In some embodiments, the second determining module 1203 is configured to: if the memory usage indicated by the second reference information meets the conditions, determine a second execution strategy and a memory reclamation strategy based on the second reference information and the first execution strategy; if the memory usage indicated by the second reference information does not meet the conditions, determine third reference information based on the second reference information and the attention computing task in the first part of the computing task, and determine the second execution strategy and the memory reclamation strategy based on the third reference information and the first execution strategy, wherein the third reference information indicates the memory usage of the storage component after unloading the inference data of the feedforward network computing task and the attention computing task in the first part of the computing task from the storage component.
[0190] Of course, the computing device described above can also include other functional units to implement other functions involved in the processor in the above method embodiments. In practical applications, the above functions can be assigned to different functional units as needed. In addition, the computing device for performing large model inference provided in the above embodiments belongs to the same concept as the above method embodiments, and its specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0191] Figure 13 This is a schematic diagram of the structure of a storage device for performing large model inference, provided in an embodiment of this application. Figure 13 As shown, the storage device includes a storage unit 1301 and a computing unit 1302, with the storage unit 1301 and the computing unit 1302 integrated.
[0192] This storage device is used for:
[0193] At least one inference request of a large model is executed according to a first execution strategy. The first execution strategy indicates that a first part of the computation task needs to be offloaded to the storage device for execution during the execution of at least one inference request. The computation unit 1302 is used to execute the computation task offloaded to the storage device according to the inference data stored in the storage unit 1301. The inference data includes at least one of model weights and input data of the computation task.
[0194] During the execution of at least one inference request, the execution of at least one inference request continues in accordance with a second execution strategy, which indicates that a second part of the computational task to be offloaded to the storage device during the execution of at least one inference request.
[0195] In some embodiments, the computational tasks offloaded to the storage device for execution include at least one of the attention computation task of the Transformer layer in the large model and the feedforward network computation task; wherein the attention computation task includes at least one of the QKV matrix generation task, attention score computation task, normalization computation task, context vector computation task, and linear projection task.
[0196] In some embodiments, the calculation unit 1302 includes a vector unit 1303 and a matrix-vector multiplication unit 1304;
[0197] Vector unit 1303 is used to perform normalization computation tasks offloaded to storage devices based on inference data;
[0198] The matrix-vector multiplication unit 1304 is used to perform tasks other than the normalization calculation task in the computational tasks offloaded to the storage device based on inference data.
[0199] In some embodiments, the storage device is further configured to: unload inference data on the storage device according to an instruction from a memory reclamation request sent by the processor.
[0200] In some embodiments, the storage device includes a high-bandwidth memory (HBM), and the HBM and the processor are connected via a bus; wherein, the storage unit 1301 is a storage bank in the HBM, and the bank is integrated with the computing unit 1302.
[0201] Of course, the aforementioned storage device may also include other functional units to implement other functions involved in the storage components in the above method embodiments. In practical applications, the above functions can be assigned to different functional units as needed. In addition, the storage device for performing large model inference provided in the above embodiments belongs to the same concept as the above method embodiments, and its specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0202] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with substantially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. These terms are merely used to distinguish one element from another. For example, without departing from the various examples described, a first operator can be referred to as a second operator, and similarly, a second operator can be referred to as a first operator. Both the first and second operators can be operators, and in some cases, they can be separate and distinct operators.
[0203] In this application, the term "at least one" means one or more, and the term "multiple" means two or more. For example, multiple operators means two or more operators.
[0204] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0205] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of program structure information. This program structure information includes one or more program instructions. When these program instructions are loaded and executed on a computing device, the processes or functions according to the embodiments of this application are generated, in whole or in part.
[0206] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0207] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A large-scale model reasoning method, characterized in that, Applied to a heterogeneous inference system, the heterogeneous inference system including a processor and a storage component, the storage component including a storage unit and a computing unit, the storage unit being integrated with the computing unit, the method comprising: The processor determines a first execution strategy based on at least one inference request of the large model. The first execution strategy indicates a first part of the computational task that needs to be offloaded to the storage component for execution during the execution of the at least one inference request. The processor and the storage component jointly execute the at least one inference request according to the first execution strategy, wherein the computing unit executes a computing task offloaded to the storage component based on the inference data stored in the storage unit, and the inference data includes at least one of model weights and input data of the computing task; During the execution of the at least one inference request, the processor determines a second execution strategy based on the first execution strategy and the memory usage state of the storage component. The second execution strategy indicates a second part of the computational task that needs to be offloaded to the storage component during the execution of the at least one inference request. The processor and the storage component continue to jointly execute the at least one inference request in accordance with the second execution strategy.
2. The method according to claim 1, characterized in that, The computational task includes at least one of the attention computation task of the Transformer layer and the feedforward network computation task in the large model; The attention calculation task includes at least one of the following: query-key-value QKV matrix generation task, attention score calculation task, normalization calculation task, context vector calculation task, and linear projection task.
3. The method according to claim 2, characterized in that, The computation unit includes a vector unit and a matrix-vector multiplication unit; The computing unit executes computing tasks offloaded to the storage component based on the inference data stored in the storage unit, including: The vector unit performs a normalization calculation task that is offloaded to the storage component based on the inference data; The matrix-vector multiplication unit executes tasks other than the normalization calculation task from the computation tasks offloaded to the storage component based on the inference data.
4. The method according to any one of claims 1 to 3, characterized in that, The processor determines a first execution strategy based on at least one inference request from the large model, including: The processor determines the first execution strategy based on the number of the at least one inference request, the number of storage components, the number of processors, the number of storage units in the storage components, and the number of attention heads in the Transformer layer of the large model; In this context, one storage unit in the storage component corresponds to at least one attention head of the Transformer layer in the large model.
5. The method according to claim 4, characterized in that, The storage component includes multiple storage unit groups, and each storage unit group includes multiple storage units; The processor and the storage component jointly execute the at least one inference request according to the first execution strategy, including: The processor and each of the plurality of storage unit groups jointly execute multiple inference requests in accordance with the first execution strategy, wherein each storage unit group corresponds to at least one inference request.
6. The method according to any one of claims 1 to 5, characterized in that, During the execution of the at least one inference request, the processor determines a second execution strategy based on the first execution strategy and the memory usage state of the storage component, including: During the decoding phase of the at least one inference request, the processor determines the second execution strategy and the memory reclamation strategy based on the first execution strategy and the memory usage state of the storage component. The memory reclamation strategy indicates the inference data that needs to be unloaded from the storage component during the decoding phase of the at least one inference request. The method further includes: The processor sends a memory reclamation request to the storage component according to the memory reclamation strategy; The storage component unloads inference data from the storage component according to the instruction of the memory reclamation request.
7. The method according to claim 6, characterized in that, The processor determines the second execution strategy and the memory reclamation strategy based on the first execution strategy and the memory usage state of the storage component, including: The processor determines first reference information based on the first execution strategy and the memory usage status. The first reference information indicates the memory usage of the storage component when the at least one inference request is executed according to the first execution strategy. The processor determines the second execution strategy and the memory reclamation strategy based on the first reference information and the type of the first part of the computing task.
8. The method according to claim 7, characterized in that, The first part of the computational tasks includes the attention computation task of the Transformer layer and the feedforward network computation task in the large model; The processor determines the second execution strategy and the memory reclamation strategy based on the first reference information and the type of the first part of the computing task, including: The processor determines second reference information based on the first reference information and the feedforward network computing task in the first part of the computing task. The second reference information indicates the memory usage of the storage component after the inference data of the feedforward network computing task in the first part of the computing task is unloaded from the storage component. The processor determines the second execution strategy and the memory reclamation strategy based on the second reference information and the first execution strategy.
9. The method according to claim 8, characterized in that, The processor determines the second execution strategy and the memory reclamation strategy based on the second reference information and the first execution strategy, including: If the memory usage indicated by the second reference information meets the conditions, the processor determines the second execution strategy and the memory reclamation strategy based on the second reference information and the first execution strategy; If the memory usage indicated by the second reference information does not meet the conditions, the processor determines third reference information based on the second reference information and the attention computing task in the first part of the computing task, and determines the second execution strategy and the memory reclamation strategy based on the third reference information and the first execution strategy, wherein the third reference information indicates the memory usage of the storage component after unloading the inference data of the feedforward network computing task and the attention computing task in the first part of the computing task from the storage component.
10. The method according to any one of claims 1 to 9, characterized in that, The processor includes a central processing unit (CPU) and an accelerator. The CPU is used to control the accelerator, and the CPU and the accelerator are connected via a bus. The processor and the storage component jointly execute the at least one inference request according to the first execution strategy, including: the CPU controlling the accelerator to execute tasks other than the computing task in the at least one inference request according to the first execution strategy and controlling the accelerator to offload the computing task to the storage component for execution.
11. The method according to any one of claims 1 to 10, characterized in that, The storage component is a high-bandwidth memory (HBM), and the storage component and the processor are connected via a bus. The storage unit is a storage bank in the HBM, and the bank is integrated with the computing unit.
12. A computing device for performing large model inference, characterized in that, The computing device and the storage component are connected. The storage component includes a storage unit and a computing unit, and the storage unit is integrated with the computing unit. The computing device includes: The first determining module is configured to determine a first execution strategy based on at least one inference request of the large model, wherein the first execution strategy indicates a first part of the computational task that needs to be offloaded to the storage component for execution during the execution of the at least one inference request; An execution module is configured to execute, in accordance with the first execution strategy, the computational tasks in the at least one inference request that were not offloaded to the storage component; The second determining module is used to determine a second execution strategy based on the first execution strategy and the memory usage state of the storage component during the execution of the at least one inference request. The second execution strategy indicates a second part of the computational task that needs to be offloaded to the storage component for execution during the execution of the at least one inference request. The execution module is further configured to continue executing the computational tasks in the at least one inference request that were not offloaded to the storage component in accordance with the second execution strategy.
13. The computing device according to claim 12, characterized in that, The computational tasks that need to be offloaded to the storage component for execution include at least one of the attention computation tasks of the Transformer layer and the feedforward network computation tasks in the large model; The attention calculation task includes at least one of the following: query-key-value QKV matrix generation task, attention score calculation task, normalization calculation task, context vector calculation task, and linear projection task.
14. The computing device according to claim 12 or 13, characterized in that, The first determining module is used for: The first execution strategy is determined based on the number of at least one inference request, the number of storage components, the number of processors, the number of storage units in the storage components, and the number of attention heads in the Transformer layer of the large model; In this context, one storage unit in the storage component corresponds to at least one attention head of the Transformer layer in the large model.
15. The computing device according to any one of claims 12 to 14, characterized in that, The second determining module is used for: During the decoding phase of the at least one inference request, a second execution strategy and a memory reclamation strategy are determined based on the first execution strategy and the memory usage status of the storage component. The memory reclamation strategy indicates the inference data that needs to be unloaded from the storage component during the decoding phase of the at least one inference request. The computing device further includes a sending module, configured to: send a memory reclamation request to the storage component according to the memory reclamation strategy.
16. The computing device according to claim 15, characterized in that, The second determining module is used for: Based on the first execution strategy and the memory usage status, first reference information is determined, and the first reference information indicates the memory usage of the storage component when the at least one inference request is continued to be executed in accordance with the first execution strategy. Based on the first reference information and the type of the first part of the computational task, the second execution strategy and the memory reclamation strategy are determined.
17. The computing device according to claim 16, characterized in that, The first part of the computational tasks includes the attention computation task of the Transformer layer and the feedforward network computation task in the large model; The second determining module is used for: Based on the first reference information and the feedforward network computing task in the first part of the computing task, a second reference information is determined, the second reference information indicating the memory usage of the storage component after the inference data of the feedforward network computing task in the first part of the computing task is unloaded from the storage component. Based on the second reference information and the first execution strategy, the second execution strategy and the memory reclamation strategy are determined.
18. The computing device according to claim 17, characterized in that, The second determining module is used for: If the memory usage indicated by the second reference information meets the conditions, the second execution strategy and the memory reclamation strategy are determined based on the second reference information and the first execution strategy. If the memory usage indicated by the second reference information does not meet the conditions, a third reference information is determined based on the second reference information and the attention computing task in the first part of the computing task, and a second execution strategy and a memory reclamation strategy are determined based on the third reference information and the first execution strategy, wherein the third reference information indicates the memory usage of the storage component after unloading the inference data of the feedforward network computing task and the attention computing task in the first part of the computing task from the storage component.
19. A storage device for performing large model inference, characterized in that, The storage device is connected to the processor, and the storage device includes a storage unit and a computing unit, wherein the storage unit is integrated with the computing unit. The storage device is used for: At least one inference request of a large model is executed according to a first execution strategy, wherein the first execution strategy indicates that a first part of the computational task needs to be offloaded to the storage device for execution during the execution of the at least one inference request, wherein the computing unit is used to execute the computational task offloaded to the storage device according to the inference data stored in the storage unit, wherein the inference data includes at least one of model weights and input data of the computational task; During the execution of the at least one inference request, the at least one inference request continues to be executed in accordance with a second execution strategy, the second execution strategy indicating that a second part of the computational task to be offloaded to the storage device during the execution of the at least one inference request.
20. The storage device according to claim 19, characterized in that, The computational task includes at least one of the attention computation task of the Transformer layer and the feedforward network computation task in the large model; The attention calculation task includes at least one of the following: query-key-value QKV matrix generation task, attention score calculation task, normalization calculation task, context vector calculation task, and linear projection task.
21. The storage device according to claim 20, characterized in that, The computation unit includes a vector unit and a matrix-vector multiplication unit; The vector unit is used to perform a normalization calculation task that is offloaded to the storage device based on the inference data; The matrix-vector multiplication unit is used to execute tasks other than the normalization calculation task in the computing tasks that are offloaded to the storage device, based on the inference data.
22. The storage device according to any one of claims 19 to 21, characterized in that, The storage device is also used for: Indication data on the storage device is unloaded according to the memory reclamation request sent by the processor.
23. The storage device according to any one of claims 19 to 22, wherein the storage device includes high-bandwidth memory (HBM), and the HBM and the processor are connected via a bus; in, The storage unit is the storage bank in the HBM, and the bank is integrated with the computing unit.
24. A heterogeneous reasoning system, characterized in that, The heterogeneous inference system includes a computing device as described in any one of claims 12 to 18 and a storage device as described in any one of claims 19 to 23.
25. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one piece of program code for implementing the large model inference method as described in any one of claims 1 to 11.
26. A computer program product, characterized in that, The computer program product is used to implement the large model inference method as described in any one of claims 1 to 11.