Big language model reasoning method and device
By offloading the self-attention computing task to the near-memory computing module, the problem of processor memory limitation is solved, the processor utilization of large language models is improved and the inference cost is reduced, especially the efficiency when processing ultra-long text.
Patent Information
- Application Number
- CN202410223757.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-18
- Filing Date
- 2024-02-28
- Publication Date
- 2025-07-18
AI Technical Summary
In the inference process of large language models, the processor cannot fully utilize its computing performance due to memory limitations, especially when processing ultra-long text, resulting in low processor utilization and high inference cost.
The bandwidth-intensive self-attention computing task is offloaded to the near-storage computing module for execution. The processor performs calculation-intensive tasks, and improves the processor utilization rate and reduces storage requirements through a parallel inference system.
It improves the utilization rate of the processor, reduces the inference cost of large language models, and improves the efficiency of processing ultra-long text sequences.
Smart Images

Figure CN120338089A_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with an application number of 202410077193.4 and an invention title of "A Method, Apparatus and Other Devices for Data Processing" filed with the National Intellectual Property Administration on January 18, 2024, the entire content of which is incorporated herein by reference. Technical Field
[0002] Embodiments of this application relate to the storage field, and in particular, to a method and apparatus for inferring a large language model. Background Art
[0003] With the development of related technologies such as machine learning, large language models (LLMs) based on machine learning technologies have been widely applied. For example, large language models can automatically generate language texts or generate responses, and can be used for tasks such as machine translation, speech recognition, question-and-answer systems, and dialogue generation.
[0004] In the current inference process of large language models, a computing device implements the inference process of the large language model through a processor, such as a graphics processing unit (GPU) and a neural processing unit (NPU). When the computing device performs inference on the large language model through the processor, it is necessary to use the high-bandwidth memory of the processor to store the intermediate data and calculation results of the inference process.
[0005] However, when the input text of the large language model is some extremely long texts, due to the memory limitation of the processor in the computing device, the computing performance of the processor cannot be fully utilized, resulting in low utilization rate of the processor. At the same time, if the computing device increases the high-bandwidth memory of the processor to process extremely long input texts, the inference cost of the large language model will increase. Summary of the Invention
[0006] Embodiments of this application provide a method for inferring a large language model. The computing device offloads the bandwidth-intensive self-attention calculation task to the near-memory computing module for execution, so that the processor of the computing device can avoid memory limitations and execute calculation-intensive tasks, thereby improving the utilization rate of the processor and reducing the inference cost of the large language model. Embodiments of this application also provide an inference apparatus for the large language model corresponding to the method for inferring the large language model, a computing device, a computing device cluster, a computer-readable storage medium, and a computer program product.
[0007] In a first aspect, an embodiment of the present application provides a method for inferring a large language model. This method is applied to a parallel inference system, which includes a processor and a near-memory computing module. Both the processor and the near-memory computing module are used to execute tasks assigned by the parallel inference system. The method provided by the first aspect includes: The parallel inference system receives an inference task of the large language model, and the inference task includes a computationally intensive task and a self-attention calculation task. The processor executes the computationally intensive task, generates intermediate data, and sends the intermediate data to the near-memory computing module. The computationally intensive task includes one or more of the following tasks of the large language model: feed-forward neural network calculation task, projection task, and layer normalization task. The near-memory computing module executes the self-attention calculation task based on the intermediate data and generates a near-memory calculation result. The processor generates an inference result corresponding to the inference task based on the near-memory calculation result.
[0008] The processor of the parallel inference system in the embodiment of the present application can send the intermediate data generated by executing the computationally intensive task to the near-memory computing module, enabling the near-memory computing module to execute the bandwidth-intensive self-attention calculation task based on this intermediate data. Thus, the self-attention calculation task is offloaded from the processor to the near-memory computing module, allowing the processor of the computing device to avoid memory limitations when executing computationally intensive tasks, improving the utilization rate of the processor. At the same time, it also reduces the demand for high-bandwidth memory in the processor and lowers the inference cost of the large language model.
[0009] In a possible implementation, the storage capacity of the near-memory computing module is greater than that of the processor, and the storage cost of the near-memory computing module is lower than that of the processor. The computing power of the processor is greater than that of the near-memory computing module. The parallel inference system can allocate the self-attention calculation task with high storage resource occupancy to the near-memory computing module for execution, and allocate the computationally intensive task with high computing power requirements to the processor for execution.
[0010] In the embodiment of the present application, since the storage capacity of the near-memory computing module is greater than that of the processor, the parallel inference system can allocate the self-attention calculation task with high storage resource occupancy to the near-memory computing module for execution, thereby reducing the processor memory limitation. At the same time, the computing power of the processor is greater than that of the near-memory computing module, and the parallel inference system can also allocate the computationally intensive task to the processor for execution, thereby improving the utilization rate of the processor and further enhancing the inference efficiency of the large language model.
[0011] In a possible implementation, one or more segments of processed text corresponding to the inference task, and the length of the processed text is greater than or equal to a first threshold. The first threshold is, for example, 1000 characters, that is, the large language model can be used to process ultra-long text sequences, and the large language model consumes more storage resources when processing ultra-long text sequences.
[0012] In the embodiments of the present application, the large language model can be used to process ultra-long text sequences. At the same time, the parallel inference system offloads the self-attention calculation tasks with high storage resource occupancy to the near-memory computing module for execution, improving the utilization rate of the processor, reducing the memory consumption of the processor, and further improving the inference efficiency of the large language model.
[0013] In a possible implementation, the parallel inference system includes multiple near-memory computing modules. When the processor sends the intermediate data to the near-memory computing modules, the processor segments the intermediate data to obtain multiple segmented intermediate data. The processor sends the multiple segmented intermediate data to the multiple near-memory computing modules. When the near-memory computing modules execute the self-attention calculation tasks based on the intermediate data, the multiple near-memory computing modules execute the self-attention calculation tasks in parallel based on the multiple segmented intermediate data to generate near-memory calculation results corresponding to the multiple near-memory computing modules. When the processor generates the inference result corresponding to the inference task based on the near-memory calculation results, the processor performs a reduction calculation on the near-memory calculation results corresponding to the multiple near-memory computing modules and generates the inference result corresponding to the inference task.
[0014] In the embodiments of the present application, the parallel inference system includes multiple near-memory computing modules. The processor can segment the intermediate data generated by executing the computationally intensive tasks, and the multiple near-memory computing modules execute the self-attention calculation tasks in parallel based on the segmented intermediate data, thereby improving the inference efficiency of the large language model. At the same time, the multiple near-memory computing modules executing the self-attention calculation tasks in parallel also reduce the hardware requirements for the near-memory computing modules, further reducing the inference cost of the large language model.
[0015] In a possible implementation, when the processor sends the multiple segmented intermediate data to the multiple near-memory computing modules, the processor sequentially and cyclically sends the multiple segmented intermediate data to the multiple near-memory computing modules so that the difference between the total lengths of the segmented intermediate data in each near-memory computing module is less than or equal to a second threshold, that is, the processor evenly distributes the segmented intermediate data to be processed by the multiple near-memory computing modules.
[0016] In the embodiments of the present application, the processor of the parallel inference system can evenly distribute the generated intermediate data to different near-memory computing modules based on the data lengths of different near-memory computing modules, thereby improving the balance of the self-attention calculation tasks executed by the multiple near-memory computing modules in the parallel inference system.
[0017] In a possible implementation, the intermediate data includes a query word matrix and key-value cache data. When the near-memory computing module performs the self-attention computing task based on the intermediate data, the near-memory computing module performs batch self-attention computing based on the query word matrix and the key-value cache data to generate a near-memory computing result. If the parallel inference system includes multiple near-memory computing modules, the multiple near-memory computing modules can perform self-attention computing in parallel to generate their respective near-memory computing results. The multiple near-memory computing modules send their respective near-memory computing results to the processor.
[0018] In the embodiment of the present application, the near-memory computing module of the parallel inference system can perform batch self-attention computing based on the query word matrix and the key-value cache data to generate a near-memory computing result, thereby improving the feasibility of the solution.
[0019] In a possible implementation, when the parallel inference system includes multiple near-memory computing modules, for the query word matrix in the intermediate data, the processor can send it to the multiple near-memory computing modules by means of broadcasting, and for the key-value cache data in the intermediate data, the processor can send it to the multiple near-memory computing modules sequentially by means of append writing.
[0020] In the embodiment of the present application, for different types of intermediate data, the processor of the parallel inference system can be sent to the near-memory computing module in different ways, thereby improving the richness of the processor sending intermediate data to the near-memory computing module.
[0021] In a possible implementation, the processor and the near-memory computing module can perform the inference task of the large language model in a pipelined and parallel manner. The large language model includes multiple transmission layers, and there are computationally intensive tasks and self-attention computing tasks in each of the multiple transmission layers. Moreover, the execution of the computationally intensive task in the current transmission layer depends on the near-memory computing result of the previous transmission layer. The parallel inference system executes the computationally intensive task and the self-attention computing task in parallel in each transmission layer. The computationally intensive task is executed by the processor, and the self-attention computing task is executed by the near-memory computing module.
[0022] In the process of the parallel inference system in the embodiment of the present application executing the inference task of the large language model, the computationally intensive task and the self-attention computing task can be executed in a pipelined and parallel manner by the processor and the near-memory computing module respectively, thereby reducing the idle time of the processor and the near-memory computing module in the parallel inference system, improving the inference efficiency of the large language model, and the full-load operation of the processor and the near-memory computing module also reduces the inference cost of the large language model.
[0023] In a possible implementation, the processor includes one or more of the following: a graphics processing unit (GPU), a neural network processing unit (NPU), and a data processing unit (DPU).
[0024] In the embodiments of the present application, the processors in the parallel inference system can be various types of processors, and different types of processors are applicable to different types of inference tasks, thereby improving the applicability of the inference method.
[0025] In a second aspect, an inference device for a large language model provided by an embodiment of the present application is applied to a parallel inference system. The parallel inference system includes a processor and a near-memory computing module. Both the processor and the near-memory computing module are used to execute the tasks assigned by the parallel inference system. The task processing device includes a transceiver unit and a processing unit. Among them, the transceiver unit is used to receive the inference task of the large language model, and the inference task includes a computationally intensive task and a self-attention calculation task. The processing unit is used to execute the computationally intensive task on the processor, generate intermediate data, and send the intermediate data to the near-memory computing module. The computationally intensive task includes one or more of the following tasks of the large language model: feed-forward neural network calculation task, projection task, and layer normalization task. The processing unit is also used to execute the self-attention calculation task on the near-memory computing module according to the intermediate data to generate a near-memory calculation result. The processing unit is also used to generate an inference result corresponding to the inference task based on the near-memory calculation result.
[0026] In a possible implementation, the storage capacity of the near-memory computing module is greater than the storage capacity of the processor, and the computing power of the processor is greater than the computing power of the near-memory computing module.
[0027] In a possible implementation, one or more segments of processed text corresponding to the inference task, and the length of the processed text is greater than or equal to the first threshold.
[0028] In a possible implementation, the parallel inference system includes multiple near-memory computing modules. The processing unit is used to segment the intermediate data on the processor to obtain multiple segmented intermediate data, and send the multiple segmented intermediate data to the multiple near-memory computing modules on the processor. Specifically, the processing unit is used to execute the self-attention calculation task in parallel on the near-memory computing module according to the multiple segmented intermediate data to generate near-memory calculation results corresponding to the multiple near-memory computing modules. Specifically, the processing unit is used to perform reduction calculation on the near-memory calculation results corresponding to the multiple near-memory computing modules on the processor and generate an inference result corresponding to the inference task.
[0029] In a possible implementation, the transceiver unit is specifically used to send multiple segmented intermediate data to multiple near-memory computing modules respectively in a loop on the processor, so that the difference between the total lengths of the segmented intermediate data in each near-memory computing module is less than or equal to the second threshold.
[0030] In a possible implementation, the intermediate data includes a query word matrix and key-value cache data. The processing unit is specifically used to perform batch self-attention calculation on the near-memory computing module based on the query word matrix and the key-value cache data to generate a near-memory calculation result.
[0031] In a possible implementation, the large language model includes multiple transmission layers. In each of the multiple transmission layers, there are compute-intensive tasks and self-attention computing tasks. Moreover, the execution of compute-intensive tasks in the current transmission layer depends on the near-memory computing results of the previous transmission layer. The processing unit is further configured to execute the compute-intensive tasks and the self-attention computing tasks in parallel in each transmission layer. The compute-intensive tasks are executed by the processor, and the self-attention computing tasks are executed by the near-memory computing module.
[0032] In a possible implementation, the processor includes one or more of the following: a graphics processing unit (GPU), a neural network processing unit (NPU), and a data processing unit (DPU).
[0033] In a third aspect, an embodiment of the present application provides a computing device. The computing device includes a processor coupled to a memory. The processor is configured to store instructions that, when executed by the processor, cause the computing device to execute the method described in the first aspect or any possible implementation of the first aspect.
[0034] In a fourth aspect, an embodiment of the present application provides a computing device cluster. The computing device cluster includes one or more computing devices. Each computing device includes a processor coupled to a memory. The processor is configured to store instructions that, when executed by the processor, cause the computing device cluster to execute the method described in the first aspect or any possible implementation of the first aspect.
[0035] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium having instructions stored thereon that, when executed, cause a computer to execute the method described in the first aspect or any possible implementation of the first aspect.
[0036] In a sixth aspect, an embodiment of the present application provides a computer program product that includes instructions that, when executed, cause a computer to implement the method described in the first aspect or any possible implementation of the first aspect.
[0037] It can be understood that the beneficial effects that can be achieved by any of the above-provided inference devices, computing devices, computing device clusters, computer-readable media, or computer program products of the large language model can refer to the beneficial effects in the corresponding method, which will not be elaborated here. Description of the Drawings
[0038] Figure 1 It is a schematic diagram of the system architecture of a large language model inference system provided by an embodiment of the present application;
[0039] Figure 2 It is a schematic flowchart of a large language model inference method provided by an embodiment of the present application;
[0040] Figure 3 Flow diagram of another large language model inference method provided by an embodiment of the present application;
[0041] Figure 4 Flow diagram of another inference method of a large language model provided by an embodiment of the present application;
[0042] Figure 5 Flow diagram of parallel inference of a processor and a near-memory computing module provided by an embodiment of the present application;
[0043] Figure 6 Flow diagram of another parallel inference of a processor and a near-memory computing module provided by an embodiment of the present application;
[0044] Figure 7 Structural diagram of an inference device of a large language model provided by an embodiment of the present application;
[0045] Figure 8 Structural diagram of a computing device provided by an embodiment of the present application;
[0046] Figure 9 Structural diagram of a computing device cluster provided by an embodiment of the present application;
[0047] Figure 10 Structural diagram of another computing device cluster provided by an embodiment of the present application. Detailed implementation manners
[0048] Embodiments of the present application provide an inference method and device for a large language model, which are used to improve the utilization rate of processors in a parallel inference system, thereby improving the inference efficiency of the large language model and reducing the inference cost of the large language model.
[0049] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0050] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0051] First, some terms involved in the embodiments of the present application are introduced to facilitate those skilled in the art to understand the technical solutions.
[0052] A large language model (LLM) is a natural language processing model based on deep learning technology. It is trained using a large amount of text data and has the ability to understand, generate, and process natural language. The large language model treats natural language text as a kind of sequence data, such as a word sequence or a character sequence, and models the statistical laws and potential semantic information of these sequence data through a deep learning model. The large language model can perform a wide range of tasks, including text summarization, translation, sentiment analysis, question answering, dialogue, etc.
[0053] A processor (x-processing unit, XPU) can also be referred to as an accelerator. In the present application, the processor refers to processors such as a graphics processing unit (GPU), a neural network processing unit (NPU), and a data processing unit (DPU).
[0054] High bandwidth memory (HBM) refers to the collective term for high bandwidth memory on the processor XPU. For example, it represents the video memory on the graphics processing unit GPU. The neural network processing unit NPU also has a corresponding high bandwidth storage component. The cost of the high bandwidth memory in the processor is relatively high.
[0055] In order to make the technical solutions of the present application clearer and easier to understand, the system architecture of the present application is introduced below with reference to the accompanying drawings.
[0056] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the system architecture of a parallel inference system provided by the embodiments of the present application. In Figure 1In the shown system architecture, the parallel inference system 10 is used to execute the inference tasks of the large language model. The parallel inference system 10 includes a processor 101 and a near-memory computing module 102. The processor 101 and the near-memory computing module 102 are connected by a high-speed transmission bus. Among them, the near-memory computing module 102 includes a processor core 1021, a memory 1022, and a flash memory 1023. The following introduces the specific functions of each part of the large language model inference system 10.
[0057] The processor 101 is used to execute the computationally intensive tasks of the large language model. The computationally intensive tasks refer to the tasks that consume relatively high computing resources of the processor during the inference process of the large language model, such as feed-forward neural network computing tasks, projection tasks, and layer normalization tasks, etc. The processor 101 is also used to send the intermediate data generated by executing the computationally intensive tasks to the near-memory computing module 102.
[0058] It should be noted that the processor 101 can be various types of processors. For example, the processor 101 can be a graphics processing unit GPU, a neural network processing unit NPU, or a data processing unit DPU, and there is no specific limitation.
[0059] The near-memory computing module 102 is used to receive the intermediate data sent by the processor 101 and execute the self-attention computing task based on the intermediate data to generate the near-memory computing result. Among them, the self-attention computing task is a bandwidth-intensive task. The bandwidth-intensive task refers to the task that consumes relatively high storage resources. The bandwidth-intensive task often involves a large amount of data reading and writing operations and requires occupying storage space. The near-memory computing module 102 is also used to send the near-memory computing result to the processor 101.
[0060] The near-memory computing module 102 includes a processor core 1021, a memory 1022, and a flash memory 1023. Among them, the processor core 1021 is used to provide the computing power for executing the self-attention computing task. The memory 1022 is used to cache the temporary data during the process of the processor core 1021 executing the self-attention computing task. The flash memory 1023 is used to store the intermediate data sent by the processor 101 and the near-memory computing result generated by the processor core 1021 executing the self-attention computing task.
[0061] It should be noted that the processor 101 and the near-memory computing module 102 in the parallel inference system 10 can be deployed on the same computing device or can be respectively deployed on different computing devices, and there is no specific limitation.
[0062] In addition, the parallel inference system 10 can be a large language model inference system composed of multiple processors 101 and multiple near-memory computing modules 102. When the processors 101 and the near-memory computing modules 102 are deployed in different computing devices, multiple processors 101 or multiple near-memory computing modules 102 in the parallel inference system 10 can also be respectively deployed in different computing devices.
[0063] It can be understood that in the embodiments of the present application, the parallel inference system 10 can be applied to the inference scenario of long texts. When the input text of the large language model is long, a large amount of intermediate data will be generated during the inference process of the large language model. If these intermediate data are stored in the high-bandwidth memory of the processor 101, it will increase the inference cost. Therefore, the parallel inference system 10 can store this intermediate data based on the near-memory computing module 102 and execute bandwidth-intensive tasks, reducing the inference cost of the long text inference scenario.
[0064] Based on Figure 1 the parallel inference system 10 shown, the present application also provides an inference method for a large language model. The inference method for the large language model provided by the embodiments of the present application will be introduced below in conjunction with the embodiments.
[0065] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an inference method for a large language model provided by an embodiment of the present application. In Figure 2 the example shown, the method includes the following steps:
[0066] Step 201. The parallel inference system receives an inference task of the large language model, and the inference task includes a computation-intensive task and a self-attention computation task.
[0067] The parallel inference system 10 receives an inference task of the large language model, where the inference task includes a computation-intensive task and a self-attention computation task. During the process of executing the inference task of the large language model, the parallel inference system 10 allocates the computation-intensive task to the processor 101 for execution, and allocates the self-attention computation task to the near-memory computing module 102 for execution.
[0068] It can be understood that since the storage capacity of the near-memory computing module 102 is greater than the storage capacity of the processor 101, and the computing power of the processor 101 is greater than the computing power of the near-memory computing module 102, the parallel inference system 10 allocates the self-attention computation task with high storage resource occupancy to the near-memory computing module 102 for execution, thereby improving the utilization rate of the processor 101.
[0069] In an example of step 201, the parallel inference system 10 can receive inference tasks of a large language model. For example, the parallel inference system 10 receives the input text of the user, and infers the output text based on the input text of the user. Among them, the input text can be a sentence or paragraph of the user's question, and the output text can be the question answer, text summary, translation result, etc. corresponding to the input text. Specifically, in the process of the parallel inference system 10 generating the output text corresponding to the input text, the parallel inference system 10 infers the input text of the user based on the multi-layer structure of the large language model. Among them, in each transmission layer of the large language model, it is necessary to perform computationally intensive tasks and self-attention calculation tasks.
[0070] Please refer to Figure 3 , Figure 3 which is a schematic diagram of an inference method for a large language model provided by an embodiment of the present application. In Figure 3 the example shown, the parallel inference system 10 receives the input text input by the user into the large language model, and performs an inference task based on the multi-layer structure of the large language model. In the incremental inference (decode) stage of the inference task, the parallel inference system 10 performs cyclic iterative inference in an autoregressive manner in multiple transmission layers of the large language model. Among them, in each transmission layer of the large language model, the processor 101 processes computationally intensive tasks respectively, and the near-memory computing module 102 performs self-attention calculation tasks.
[0071] It should be noted that the large language model in the embodiment of the present application can be used to process ultra-long text sequences, that is, the inference task of the large language model can correspond to one or more segments of processed text, and the length of the processed text is greater than or equal to the first threshold, and the first threshold is, for example, 1000 characters.
[0072] Since the large language model needs to consume more storage resources when processing ultra-long text sequences, due to the limitation of high-bandwidth storage, the length of the text sequence processed in batches by the processor 101 is limited, resulting in a relatively low utilization rate of the processor 101. Therefore, the parallel inference system 10 offloads the self-attention calculation task with high storage resource occupancy to the near-memory computing module 102 for execution, and the processor 101 and the near-memory computing module 102 process the inference task of the large language model in parallel, thereby improving the utilization rate of the processor 101.
[0073] Step 202. The processor executes computationally intensive tasks, generates intermediate data, and sends the intermediate data to the near-memory computing module.
[0074] After receiving the computationally intensive tasks allocated by the parallel inference system, the processor 101 executes the computationally intensive tasks, generates intermediate data, and sends the intermediate data to the near-memory computing module 102. Among them, the computationally intensive tasks include one or more of the following tasks of the large language model: feed-forward network (FFN) computing tasks, projection tasks, and layer normalization tasks.
[0075] Among them, the feed-forward network computing task includes the computing tasks executed by the feed-forward network during the inference process of the large language model, including tasks such as feature extraction and non-linear transformation. In the large language model, the feed-forward network usually consists of multiple fully connected layers and activation functions, and is used to perform non-linear mapping and feature extraction on the input data. For example, the feed-forward network can perform non-linear transformation and feature extraction on the input word vectors.
[0076] The projection task refers to the task of mapping the input low-dimensional word vector representation to a high-dimensional space. In the large language model, the projection operation converts the word vector representation of the model into the word vector representation of the output space, so as to further generate the output text. That is, through the projection operation, the model can map the hidden representation to the appropriate output space to generate results that meet the task requirements.
[0077] The layer normalization task is a processing task that normalizes the output in each layer of the large language model, which can accelerate the convergence of the large language model and improve the generalization ability of the large language model. In the large language model, the layer normalization operation can be used as a standard operation for each layer, and it is necessary to perform normalization calculations on the output of each layer.
[0078] In the embodiment of this application, the intermediate data generated by the processor 101 executing the computationally intensive tasks includes a query word matrix and key-value cache data. Among them, the query word matrix is, for example, a matrix composed of query vectors, and the key-value cache data is, for example, key vectors and value vectors.
[0079] Please continue to refer to Figure 3 , in Figure 3 In the example shown, during the process of the parallel inference system 10 executing the inference task of the large language model, in each transmission layer of the large language model, the processor 101 executes computationally intensive tasks such as layer normalization tasks and projection tasks, generates intermediate data, and the intermediate data is the query word matrix Q, the key vector K, and the value vector V.
[0080] In a possible implementation, after the processor 101 executes a compute-intensive task to generate intermediate data, the processor 101 segments the intermediate data to generate multiple segmented intermediate data, and sends the multiple segmented intermediate data to the near-memory computing module 102 for self-attention calculation.
[0081] Please refer to Figure 4 , Figure 4 which is a schematic diagram of an inference method for a large language model provided by an embodiment of this application. In Figure 4 the example shown, the compute-intensive tasks and bandwidth-intensive tasks of the large language model are executed in the processor 101 and the near-memory computing module 102 respectively. In each round of inference calculation of the large language model, the processor 101 needs to send intermediate data to the near-memory computing module 102 and receive the near-memory computing results sent by the near-memory computing module 102.
[0082] For example, in Figure 4 the example shown, in the inference process of each transmission layer, the processor 101 executes a projection task and a feed-forward network calculation task based on compute-intensive operators to generate intermediate data for each layer of inference, and sends the intermediate data to the near-memory computing module 102 in the inference process of each layer. The near-memory computing module 102 executes a self-attention calculation task based on bandwidth-intensive operators to generate near-memory computing results for each layer of inference, and sends the near-memory computing results to the processor 101.
[0083] In Figure 4 the example shown, the processor 101 can also segment the generated intermediate data and send the segmented intermediate data to the near-memory computing module. For example, the processor 101 segments the key-value cache generated by executing compute-intensive tasks such as layer normalization tasks and projection tasks into key-value cache data KV1, key-value cache data KV2, and key-value cache data KV2, and sends the key-value cache data KV1, key-value cache data KV2, and key-value cache data KV2 to the near-memory computing module 102 for self-attention calculation.
[0084] In a possible implementation, if the parallel inference system 10 includes multiple near-memory computing modules 102, during the process of the processor 101 sending the segmented intermediate data to the multiple near-memory computing modules 102, the processor 101 respectively and circularly sends the multiple segmented intermediate data to the multiple near-memory computing modules 102, where the difference between the total lengths of the segmented intermediate data in the multiple near-memory computing modules 102 is less than or equal to a second threshold, that is, the segmented intermediate data of the processor 101 is evenly distributed for processing by the multiple near-memory computing modules 102.
[0085] It should be noted that if the parallel inference system 10 includes multiple near-memory computing modules 102, for the query word matrix in the intermediate data, the processor 101 can send it to the multiple near-memory computing modules 102 in a broadcast manner. For the key-value cache data in the intermediate data, the processor 101 can send it to the multiple near-memory computing modules 102 sequentially in an append-write manner.
[0086] Please continue to refer to Figure 3 , in Figure 3 In the example shown, the parallel inference system 10 includes 2 near-memory computing modules, namely the near-memory computing module 1 and the near-memory computing module 2. The processor 101 executes a compute-intensive task to generate a query word matrix Q and key-value cache data. The key-value cache data includes a key vector K and a value vector V. The processor 101 sends the query word matrix Q and the key-value cache data to the near-memory computing module 1 and the near-memory computing module 2 respectively.
[0087] In Figure 3 In the example shown, in the prefill inference stage of the large language model, the processor 101 has stored some key-value cache data segments in the near-memory computing module 1 and the near-memory computing module 2. For example, the near-memory computing module 1 stores the key-value segment KV segment0, and the near-memory computing module 2 stores the key-value segment KV segment1. In the incremental inference (decode) stage of the large language model, the processor 101 sends the newly added key-value cache data to the near-memory computing module 1 and the near-memory computing module 2 in an append-write manner. To prevent the computational load imbalance between the two near-memory computing modules, the processor 101 segments the key-value cache data and sends the key-value cache data segments to the near-memory computing module 1 and the near-memory computing module 2 in a cyclic manner. This cyclic sending method is also called round-robin allocation.
[0088] For example, in Figure 3 In the example shown, when the processor 101 sends KV segment3 to the near-memory computing module 2, if the total length of the key-value cache data in the near-memory computing module 2 exceeds the preset value, the processor 101 stops sending the key-value cache data to the computing module 2 and resumes sending KV segment4 to the near-memory computing module 1, so that the difference in the total length of the key-value cache data segments between the near-memory computing module 1 and the near-memory computing module 2 is less than the second threshold.
[0089] In Figure 3 In the example shown, for the query word matrix Q generated by the processor 101 executing a compute-intensive task, the processor 101 sends the query word matrix Q to the near-memory computing module 1 and the near-memory computing module 2 in a broadcast manner.
[0090] Step 203. The near-memory computing module performs a self-attention calculation task based on the intermediate data to generate a near-memory computing result.
[0091] After receiving the intermediate data sent by the processor 101, the near-memory computing module 102 performs a self-attention calculation task based on the intermediate data to generate a near-memory computing result, and sends the near-memory computing result to the processor 101. Specifically, the intermediate data includes a query word matrix Q and key-value cache data KV, and the near-memory computing module 102 performs a batched self-attention calculation based on the query word matrix Q and the key-value cache data KV to generate a near-memory computing result.
[0092] It can be understood that if the parallel inference system 10 includes multiple near-memory computing modules 102, the multiple near-memory computing modules 102 can perform self-attention calculations in parallel to generate their respective near-memory computing results. The multiple near-memory computing modules 102 send their respective near-memory computing results to the processor 101.
[0093] Please continue Figure 3 , in Figure 3 In the example shown, after the near-memory computing module 1 and the near-memory computing module 2 respectively receive the query word matrix Q and the key-value cache data sent by the processor 101, the near-memory computing module 1 and the near-memory computing module 2 perform a batched self-attention calculation based on the KV segments of their respective near-memory computing modules to generate a near-memory computing result. The near-memory computing result of the near-memory computing module 1 is, for example, score1, and the near-memory computing result of the near-memory computing module 2 is, for example, score2. The near-memory computing module 1 and the near-memory computing module 2 respectively send their respective near-memory computing results to the processor 101.
[0094] Step 204. The processor generates an inference result corresponding to the inference task based on the near-memory computing result.
[0095] After receiving the near-memory computing result sent by the near-computing module 102, the processor 101 generates an inference result corresponding to the inference task based on the near-memory computing result. If the parallel inference system 10 includes multiple near-memory computing modules 102, after the processor 101 receives the near-memory computing results sent by the multiple near-computing modules 102, it performs a reduction calculation on the near-memory computing results corresponding to the multiple near-memory computing modules 102 and generates an inference result corresponding to the inference task.
[0096] Please continue to refer to 3, in Figure 3In the example shown, after the processor 101 receives the near-memory computing results score1 and score2 sent by the near-memory computing module 1 and the near-memory computing module 2, the processor 101 performs a reduction calculation and a residual sum on the near-memory computing results score1 and score2 to obtain the processed near-memory computing results. The processor 101 continues to execute the computationally intensive tasks of the next transport layer based on the processed near-memory computing results.
[0097] In a possible implementation, the processor 101 and the near-memory computing module 102 can process the inference tasks of the large language model in a pipelined manner. Specifically, since the large language model includes multiple transport layers, there are computationally intensive tasks and self-attention computing tasks in each of the multiple transport layers, and the execution of the computationally intensive tasks in the current transport layer depends on the near-memory computing results of the previous transport layer. The parallel inference system 10 executes the computationally intensive tasks and the self-attention computing tasks in parallel in each transport layer, where the computationally intensive tasks are executed by the processor 101 and the self-attention computing tasks are executed by the near-memory computing module 102.
[0098] Please refer to Figure 5 , Figure 5 which is a schematic diagram of another inference method for the large language model provided by the embodiments of this application. In Figure 5 the example shown, the processor 101 and the near-memory computing module 102 in the parallel inference system 10 respectively execute the feed-forward network computing tasks and the self-attention computing tasks in parallel, where the processor 101 needs the attention computing results of the near-memory computing module 102 in the previous transport layer to execute the feed-forward network computing tasks in the current transport layer. For example, Figure 5 when the processor 101 executes the FFN task, it needs to rely on the near-memory computing results of the near-memory computing module 102 executing the Attn task in the previous layer.
[0099] In Figure 5 the example shown, in order to reduce the idle time of the processor 101 and the near-memory computing module 102, the processor 101 and the near-memory computing module 102 respectively execute the feed-forward network computing tasks and the self-attention computing tasks in a pipelined parallel manner, that is, when the processor 101 executes the feed-forward network computing tasks corresponding to the previous transport layer, the near-memory computing module 102 executes the attention computing tasks corresponding to the next transport layer in parallel.
[0100] For example, in Figure 5In the illustrated example, while the processor 101 is executing the computationally intensive task of the inference task req1, the near-memory computing module 102 concurrently executes the self-attention computing task of the inference task req2. While the processor 101 is executing the computationally intensive task of the inference task req2, the near-memory computing module 102 concurrently executes the self-attention computing task of the next transmission layer of the inference task req1, thereby preventing the processor 101 and the near-memory computing module 102 from being idle.
[0101] Please refer to Figure 6 , Figure 6 which is a schematic diagram of pipelined parallel inference of a large language model provided by an embodiment of the present application. In Figure 6 the illustrated example, the processor 101 and the near-memory computing module 102 respectively execute the feed-forward network computing task and the self-attention computing task in parallel. Among them, the inference task rep1 and the inference task req3 are alternately processed by the processor 101 and the near-memory computing module 102. While the near-memory computing module 102 is executing the self-attention computing task corresponding to rep1, the processor 101 is concurrently executing the feed-forward network computing task corresponding to rep3.
[0102] For example, in Figure 6 the illustrated example, between the S+0 moment and the S+1 moment, while the near-memory computing module 102 is executing the self-attention computing tasks Attn1-1 and Attn1-2 and the self-attention computing tasks Attn2-1 and Attn2-2, the processor 101 is concurrently executing the feed-forward network computing tasks FNN3-1 and FNN4-1.
[0103] In Figure 6 the illustrated example, multiple inference tasks can also be executed in parallel in the processor 101. For example, the processor 101 concurrently executes the inference task rep3 and the inference task rep4 between the S+0 moment and the S+1 moment. The processor 101 concurrently executes the inference task rep1 and the inference task rep2 between the S+1 moment and the S+2 moment.
[0104] It can be seen from the above embodiments that the processor of the parallel inference system in the embodiments of the present application can send the intermediate data generated by executing the computationally intensive task to the near-memory computing module, so that the near-memory computing module executes the wide-intensive self-attention computing task based on the intermediate data, thereby realizing offloading the self-attention computing task to the near-memory computing module for execution, enabling the processor of the parallel inference system to avoid memory limitations in executing computationally intensive tasks, and thus improving the utilization rate of the processor.
[0105] Based on the above method embodiments, an embodiment of the present application further provides an inference device for a large language model. The inference device for a large language model provided by the embodiments of the present application will be specifically introduced below.
[0106] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an inference device for a large language model provided by an embodiment of the present application. In Figure 7 the example shown, the inference device 700 of the large language model is used to implement each step executed by the parallel inference system in the above embodiments. The inference device 700 of the large language model includes a transceiver unit 701 and a processing unit 702.
[0107] Among them, the transceiver unit 701 is used to receive the inference task of the large language model. The inference task includes a compute-intensive task and a self-attention calculation task. The processing unit 702 is used to execute the compute-intensive task on the processor, generate intermediate data, and send the intermediate data to the near-memory computing module. The compute-intensive task includes one or more of the following tasks of the large language model: feed-forward neural network calculation task, projection task, and layer normalization task. The processing unit 702 is also used to execute the self-attention calculation task on the near-memory computing module according to the intermediate data to generate a near-memory computing result. The processing unit 702 is also used to generate an inference result corresponding to the inference task based on the near-memory computing result.
[0108] In a possible implementation, the storage capacity of the near-memory computing module is greater than the storage capacity of the processor, and the computing power of the processor is greater than the computing power of the near-memory computing module.
[0109] In a possible implementation, one or more segments of processing text corresponding to the inference task, and the length of the processing text is greater than or equal to the first threshold.
[0110] In a possible implementation, the parallel inference system includes multiple near-memory computing modules. The processing unit 702 is used to segment the intermediate data on the processor to obtain multiple segmented intermediate data, and send the multiple segmented intermediate data to the multiple near-memory computing modules on the processor. Specifically, the processing unit 702 is used to execute the self-attention calculation task in parallel on the near-memory computing module according to the multiple segmented intermediate data to generate near-memory computing results corresponding to the multiple near-memory computing modules. Specifically, the processing unit 702 is used to perform a reduction calculation on the near-memory computing results corresponding to the multiple near-memory computing modules on the processor and generate an inference result corresponding to the inference task.
[0111] In a possible implementation, the transceiver unit 701 is specifically used to send multiple segmented intermediate data to multiple near-memory computing modules in a loop on the processor, so that the difference between the total lengths of the segmented intermediate data in each near-memory computing module is less than or equal to the second threshold.
[0112] In a possible implementation, the intermediate data includes a query word matrix and key-value cache data. The processing unit 702 is specifically configured to perform batch self-attention calculation based on the query word matrix and the key-value cache data in the near-memory computing module to generate a near-memory computing result.
[0113] In a possible implementation, the large language model includes multiple transmission layers. In each of the multiple transmission layers, there are computationally intensive tasks and self-attention calculation tasks. Moreover, the execution of the computationally intensive tasks in the current transmission layer depends on the near-memory computing results of the previous transmission layer. The processing unit 702 is further configured to execute the computationally intensive tasks and the self-attention calculation tasks in parallel in each transmission layer. The computationally intensive tasks are executed by the processor, and the self-attention calculation tasks are executed by the near-memory computing module.
[0114] In a possible implementation, the processor includes one or more of the following: a graphics processing unit (GPU), a neural network processing unit (NPU), and a data processing unit (DPU).
[0115] It can be understood that the transceiver unit 701 and the processing unit 702 in the inference device 700 of the large language model can be used as functional modules and Figure 1 there is a mapping with each module in the large language model inference system 10 in
[0116] It should be understood that the division of units in the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And the units in the device can all be implemented in the form of software called by processing elements; they can also all be implemented in hardware; or some units can be implemented in the form of software called by processing elements, and some units can be implemented in hardware. For example, each unit can be a separately established processing element, or can be integrated in a certain chip of the device. In addition, it can also be stored in the memory in the form of a program, and the function of the unit is called and executed by a certain processing element of the device. In addition, these units can be fully or partially integrated together or independently implemented. The processing element mentioned here can also be called a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above units can be implemented through the hardware integrated logic circuit in the processor element or in the form of software called by the processing element.
[0117] It is worth noting that for the above method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.
[0118] Other reasonable combinations of steps that can be conceived by those skilled in the art based on the above description also fall within the protection scope of this application. Secondly, those skilled in the art should also be familiar that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for this application.
[0119] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a computing device provided by an embodiment of this application. As Figure 8 shown, the computing device 800 includes: a processor 801, a memory 802, a communication interface 803, and a bus 804. The processor 801, the memory 802, and the communication interface 803 are coupled through a bus (not labeled in the figure). The memory 802 stores instructions. When the execution instructions in the memory 802 are executed, the computing device 800 executes the method executed by the computing device in the above method embodiment.
[0120] The computing device 800 can be one or more integrated circuits configured to implement the above method. For example: one or more application specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms. Again, when the units in the device can be implemented in the form of a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call programs. Again, these units can be integrated together to be implemented in the form of a system-on-a-chip (SOC).
[0121] The processor 801 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor.
[0122] The memory 802 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0123] The executable program code is stored in the memory 802, and the processor 801 executes the executable program code to implement the functions of the foregoing units or modules respectively, so as to implement the inference method of the above large language model. That is to say, the instructions for executing the inference method of the above large language model are stored on the memory 802.
[0124] The communication interface 803 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 800 and other devices or a communication network.
[0125] In addition to including a data bus, the bus 804 may further include a power bus, a control bus, a status signal bus, etc. The bus may be a peripheral component interconnect express (PCIe) bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. The bus may be divided into an address bus, a data bus, a control bus, etc.
[0126] Please refer to Figure 9 , Figure 9 which is a schematic diagram of a computing device cluster provided by an embodiment of the present application. As Figure 9 shown, the computing device cluster 900 includes at least one computing device 800.
[0127] As Figure 9 shown, the computing device cluster 900 includes at least one computing device 800. Instructions for executing the above-mentioned inference method of the large language model may be stored in the memories 802 of one or more of the computing devices 800 in the computing device cluster 900.
[0128] In some possible implementation manners, partial instructions for executing the above-mentioned inference method of the large language model may also be stored separately in the memories 802 of one or more of the computing devices 800 in the computing device cluster 900. In other words, a combination of one or more computing devices 800 may jointly execute the instructions for executing the above-mentioned inference method of the large language model.
[0129] It should be noted that the memories 802 in different computing devices 800 in the computing device cluster 900 may store different instructions, respectively for executing partial functions of the above-mentioned inference device of the large language model. That is, the instructions stored in the memories 802 in different computing devices 800 may implement the functions of one or more modules in the transceiver unit and the processing unit.
[0130] In some possible implementation manners, one or more of the computing devices 800 in the computing device cluster 900 may be connected through a network. Among them, the network may be a wide area network or a local area network, etc.
[0131] Please refer to Figure 10 , Figure 10The figure is a schematic diagram of the network connection of computer devices in a computer cluster provided by an embodiment of this application. As Figure 10 shown, two computing devices 800A and 800B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device.
[0132] In a possible implementation, the memory in computing device 800A stores instructions for executing the functions of the transceiver unit. At the same time, the memory in computing device 800B stores instructions for executing the functions of the processing unit.
[0133] It should be understood that Figure 10 the functions of computing device 800A shown in
[0134] In another embodiment of this application, a computer-readable storage medium is further provided. The computer-readable storage medium stores computer-executable instructions. When the processor of the device executes the computer-executable instructions, the device executes the method executed by the computing device in the above method embodiment.
[0135] In another embodiment of this application, a computer program product is further provided. The computer program product includes computer-executable instructions, and the computer-executable instructions are stored in a computer-readable storage medium. When the processor of the device executes the computer-executable instructions, the device executes the method executed by the computing device in the above method embodiment.
[0136] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0137] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0138] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0139] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0140] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.
Claims
1. An inference method for a large language model, characterized in that, Applied to a parallel inference system, the parallel inference system includes a processor and a near-memory computing module, both the processor and the near-memory computing module are used to execute the tasks assigned by the parallel inference system, and the method includes: The parallel inference system receives an inference task of a large language model, and the inference task includes a computationally intensive task and a self-attention calculation task; The processor executes the computationally intensive task, generates intermediate data, and sends the intermediate data to the near-memory computing module. The computationally intensive task includes one or more of the following tasks of the large language model: feed-forward neural network calculation task, projection task, and layer normalization task; The near-memory computing module executes the self-attention calculation task based on the intermediate data and generates a near-memory calculation result; The processor generates an inference result corresponding to the inference task based on the near-memory calculation result.
2. The method according to claim 1, wherein The storage capacity of the near-memory computing module is greater than the storage capacity of the processor, and the computing power of the processor is greater than the computing power of the near-memory computing module.
3. The method according to claim 1 or 2, characterized in that, One or more segments of processed text corresponding to the inference task, and the length of the processed text is greater than or equal to a first threshold.
4. The method according to any one of claims 1 to 3, characterized in that, The parallel inference system includes multiple near-memory computing modules, and the processor sending the intermediate data to the near-memory computing module includes: The processor segments the intermediate data to obtain multiple segmented intermediate data; The processor sends the multiple segmented intermediate data to the multiple near-memory computing modules; The near-memory computing module executing the self-attention calculation task based on the intermediate data includes: The multiple near-memory computing modules execute the self-attention calculation task in parallel based on the multiple segmented intermediate data and generate near-memory calculation results corresponding to the multiple near-memory computing modules; The processor generating an inference result corresponding to the inference task based on the near-memory calculation result includes: The processor performs a reduction calculation on the near-memory calculation results corresponding to the multiple near-memory computing modules and generates an inference result corresponding to the inference task.
5. The method according to claim 4, characterized in that, The processor sending the multiple segmented intermediate data to the multiple near-memory computing modules includes: The processor circulates and sends the multiple segmented intermediate data to the multiple near-memory computing modules respectively, so that the difference between the total lengths of the segmented intermediate data in each near-memory computing module is less than or equal to a second threshold.
6. The method according to any one of claims 1 to 3, characterized in that, The intermediate data includes a query word matrix and key-value cache data, and the near-memory computing module executing the self-attention calculation task based on the intermediate data further includes: The near-memory computing module performs batch self-attention calculation based on the query word matrix and the key-value cache data and generates the near-memory calculation result.
7. The method according to any one of claims 1 to 4, characterized in that, The large language model includes multiple transmission layers, and in each of the multiple transmission layers, there are the computationally intensive task and the self-attention calculation task. Moreover, executing the computationally intensive task in the current transmission layer depends on the near-memory calculation result of the previous transmission layer, and the method further includes: The parallel inference system executes the computationally intensive task and the self-attention calculation task in parallel at each transport layer. The computationally intensive task is executed by the processor, and the self-attention calculation task is executed by the near-memory computing module.
8. The method according to any one of claims 1 to 5, characterized in that The processor includes one or more of the following: a graphics processing unit (GPU), a neural network processing unit (NPU), and a data processing unit (DPU).
9. An inference device for a large language model, characterized in that, Applied to a parallel inference system, the parallel inference system includes a processor and a near-memory computing module. Both the processor and the near-memory computing module are used to execute the tasks assigned by the parallel inference system. The device includes: A transceiver unit for receiving the inference task of the large language model, where the inference task includes a computationally intensive task and a self-attention calculation task; A processing unit for executing the computationally intensive task in the processor to generate intermediate data and sending the intermediate data to the near-memory computing module. The computationally intensive task includes one or more of the following tasks of the large language model: a feed-forward neural network calculation task, a projection task, and a layer normalization task; The processing unit is further configured to execute the self-attention calculation task in the near-memory computing module based on the intermediate data to generate a near-memory calculation result; The processing unit is further configured to generate an inference result corresponding to the inference task based on the near-memory calculation result.
10. The device according to claim 9, characterized in that, The storage capacity of the near-memory computing module is greater than the storage capacity of the processor, and the computing power of the processor is greater than the computing power of the near-memory computing module.
11. The device according to claim 9 or 10, characterized in that, One or more segments of processed text corresponding to the inference task, where the length of the processed text is greater than or equal to a first threshold.
12. The device according to any one of claims 9 to 11, characterized in that, The parallel inference system includes multiple near-memory computing modules, and the processing unit is configured to: Segment the intermediate data in the processor to obtain multiple segmented intermediate data; Send the multiple segmented intermediate data to the multiple near-memory computing modules in the processor; Specifically, the processing unit is configured to: Execute the self-attention calculation task in parallel in the near-memory computing module based on the multiple segmented intermediate data to generate near-memory calculation results corresponding to the multiple near-memory computing modules; Specifically, the processing unit is configured to: Perform a reduction calculation on the near-memory calculation results corresponding to the multiple near-memory computing modules in the processor and generate an inference result corresponding to the inference task.
13. The device according to claim 12, characterized in that, Specifically, the transceiver unit is configured to: Send the multiple segmented intermediate data to the multiple near-memory computing modules in a loop in the processor, so that the difference between the total lengths of the segmented intermediate data in each near-memory computing module is less than or equal to a second threshold.
14. The device according to any one of claims 9 to 11, characterized in that, The intermediate data includes a query word matrix and key-value cache data. Specifically, the processing unit is configured to: Perform batch self-attention calculation in the near-memory computing module based on the query word matrix and the key-value cache data to generate the near-memory calculation result.
15. The device according to any one of claims 9 to 12, characterized in that The large language model includes multiple transmission layers. In each of the multiple transmission layers, there exist the compute-intensive tasks and the self-attention computing tasks. Moreover, the execution of the compute-intensive tasks in the current transmission layer depends on the near-memory computing results of the previous transmission layer. The processing unit is further configured to: Execute the compute-intensive tasks and the self-attention computing tasks in parallel in each of the transmission layers. The compute-intensive tasks are executed by the processor, and the self-attention computing tasks are executed by the near-memory computing module.
16. The device according to any one of claims 9 to 15, characterized in that, The processor includes one or more of the following: a graphics processing unit (GPU), a neural network processing unit (NPU), and a data processing unit (DPU).
17. A computing device, characterized in that, Comprising a processor, the processor being coupled to a memory. The processor is configured to store instructions that, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1 to 8.
18. A cluster of computing devices, characterized in that, Comprising at least one computing device, the computing device including a processor, the processor being coupled to a memory. The processor is configured to store instructions that, when executed by the processor, cause the computing device cluster to perform the method according to any one of claims 1 to 8.
19. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instructions are executed, cause the computer to perform the method according to any one of claims 1 to 8.
20. A computer program product, comprising instructions therein, characterized in that, When the instructions are executed, cause the computer to implement the method according to any one of claims 1 to 8.
Citation Information
Cited By
Task allocation method and device, equipment and medium
CN121210150A
Streaming thinking and reasoning system and method for large language model
CN121350102A
Large language model reasoning method and system based on block storage device
CN122065970A
Reasoning method and device for large language model
WO2025152398A1