Reasoning method and device for large language model

By offloading the self-attention computing task to the near-memory computing module, the problem of processor memory limitation is solved, the processor is efficiently utilized and cost-reduced, and the inference efficiency of the large language model is improved.

WO2025152398A1PCT designated stage expired Publication Date: 2025-07-24HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Patent Information

Application Number
PCT/CN2024/109747
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-28
Filing Date
2024-08-05
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

In the inference process of large language models, the processor cannot fully utilize its computing performance due to memory limitations, and increasing the processor's high bandwidth memory to process ultra-long input text will lead to an increase inference cost.

Method used

By offloading bandwidth-intensive self-attention computing tasks to the near-storage computing module for execution, the processor avoids memory limitations and improves processor utilization. The parallel processor and near-storage computing module perform pipeline parallel computing.

Benefits of technology

It improves the utilization rate of the processor, reduces the inference cost and cost of large language models, and improves the inference efficiency of processing ultra-long text sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024109747_24072025_PF_FP_ABST
    Figure CN2024109747_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in embodiments of the present application are a reasoning method and device for a large language model, which are used for reducing the bandwidth consumption of large language model reasoning. The method of the embodiments of the present application comprises: a parallel reasoning system receives a reasoning task of a large language model, wherein the reasoning task comprises a compute-intensive task and a self-attention computing task; a processor executes the compute-intensive task, generates intermediate data, and sends the intermediate data to a near-memory computing module, wherein the compute-intensive task comprises one or more tasks of the large language model: a feedforward neural network computing task, a projection task, and a layer normalization task; the near-memory computing module executes the self-attention computing task on the basis of the intermediate data to generate a near-memory computing result; and on the basis of the near-memory computing result, the processor generates a reasoning result corresponding to the reasoning task.
Need to check novelty before this filing date? Find Prior Art

Description

A large language model reasoning method and device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 18, 2024, with application number 202410077193.4 and application name “A method, device and other equipment for data processing”, and claims priority to the Chinese patent application filed with the State Intellectual Property Office on February 28, 2024, with application number 202410223757.0 and application name “A reasoning method and device for a large language model”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of storage, and in particular to a method and device for reasoning about a large language model. Background Art

[0003] With the development of machine learning and other related technologies, large language models (LLMs) based on machine learning have gained widespread application. For example, large language models can automatically generate text or responses, and can be used for tasks such as machine translation, speech recognition, question-answering systems, and dialogue generation.

[0004] Currently, computing devices use processors, such as graphics processing units (GPUs) and neural processing units (NPUs), to perform inference on large language models. These processors require high-bandwidth memory to store intermediate data and computational results.

[0005] However, when the input text for large language models is extremely long, the computing device's processor memory limitations prevent it from fully utilizing its computing performance, resulting in low processor utilization. Furthermore, increasing the processor's high-bandwidth memory to handle extremely long input text increases the cost of inference for large language models.

[0006] Summary of the Invention

[0007] Embodiments of the present application provide a method for reasoning about a large language model. By offloading bandwidth-intensive self-attention computation tasks to a near-memory computing module, a computing device enables the computing device's processor to avoid memory limitations and execute computationally intensive tasks, thereby improving processor utilization and reducing the cost of reasoning about the large language model. Embodiments of the present application also provide a large language model reasoning apparatus, computing device, computing device cluster, computer-readable storage medium, and computer program product corresponding to the large language model reasoning method.

[0008] In the first aspect, an embodiment of the present application provides an inference method for a large language model, which is applied to a parallel inference system, wherein the parallel inference system includes a processor and a near-memory computing module, and both the processor and the near-memory computing module are used to execute tasks assigned by the parallel inference system. The method provided in the first aspect includes: the parallel inference system receives an inference task of a large language model, and the inference task includes a computationally intensive task and a self-attention computing task. The processor executes the computationally intensive task, generates intermediate data, and sends the intermediate data to the near-memory computing module. The computationally intensive task includes one or more of the following tasks of the large language model: a feedforward neural network computing task, a projection task, and a layer normalization task. The near-memory computing module executes the self-attention computing task based on the intermediate data and generates a near-memory computing result. The processor generates an inference result corresponding to the inference task based on the near-memory computing result.

[0009] The processor of the parallel reasoning system of the embodiment of the present application can generate intermediate data for executing computationally intensive tasks and send it to the near-memory computing module, so that the near-memory computing module executes bandwidth-intensive self-attention computing tasks based on these intermediate data, thereby offloading the self-attention computing tasks from the processor to the near-memory computing module for execution, so that the processor of the computing device can avoid memory limitations to execute computationally intensive tasks, thereby improving the utilization of the processor. At the same time, it also reduces the demand for high-bandwidth memory in the processor and reduces the inference cost of large language models.

[0010] In one possible implementation, the near-memory computing module has a larger storage capacity than the processor, the near-memory computing module's storage cost is lower than the processor's storage cost, and the processor's computing power is greater than the near-memory computing module's. The parallel inference system can assign self-attention computation tasks that require high storage resources to the near-memory computing module and assign computationally intensive tasks that require high computing power to the processor.

[0011] In the embodiment of the present application, since the storage capacity of the near-memory computing module is greater than the storage capacity of the processor, the parallel reasoning system can assign self-attention computing tasks that occupy a high amount of storage resources to the near-memory computing module for execution, thereby reducing the memory limitation of the processor. At the same time, the computing power of the processor is greater than the computing power of the near-memory computing module, and the parallel reasoning system can also assign computing-intensive tasks to the processor for execution, thereby improving the utilization of the processor and further improving the reasoning efficiency of the large language model.

[0012] In one possible implementation, the length of one or more processed texts corresponding to the inference task is greater than or equal to a first threshold, such as 1,000 characters. That is, a large language model can be used to process ultra-long text sequences, and a large language model needs to consume more storage resources when processing ultra-long text sequences.

[0013] In the embodiment of the present application, the large language model can be used to process ultra-long text sequences. At the same time, the parallel reasoning system offloads the self-attention calculation tasks that occupy a large amount of storage resources to the near-memory computing module for execution, thereby improving the utilization of the processor, reducing the memory consumption of the processor, and further improving the reasoning efficiency of the large language model.

[0014] In one possible implementation, the parallel reasoning system includes multiple near-memory computing modules. When the processor sends intermediate data to the near-memory computing module, the processor segments the intermediate data to obtain multiple segmented intermediate data. The processor sends the multiple segmented intermediate data to multiple near-memory computing modules. When the near-memory computing module performs a self-attention computing task based on the intermediate data, the multiple near-memory computing modules perform the self-attention computing task in parallel based on the multiple segmented intermediate data, generating near-memory computing results corresponding to the multiple near-memory computing modules. When the processor generates an inference result corresponding to an inference task based on the near-memory computing result, the processor performs an advance-reduction calculation on the near-memory computing results corresponding to the multiple near-memory computing modules, and generates an inference result corresponding to the inference task.

[0015] In the embodiment of the present application, the parallel reasoning system includes multiple near-memory computing modules. The processor can segment the intermediate data generated by executing computationally intensive tasks, and multiple near-memory computing modules can execute self-attention computing tasks in parallel based on the intermediate data segments, thereby improving the reasoning efficiency of the large language model. At the same time, the parallel execution of self-attention computing tasks by multiple near-memory computing modules also reduces the hardware requirements for the near-memory computing modules, further reducing the reasoning cost of the large language model.

[0016] In one possible implementation, when the processor sends multiple segmented intermediate data to multiple near-memory computing modules, the processor cyclically sends multiple segmented intermediate data to the multiple near-memory computing modules respectively, so that the difference between the total lengths of the segmented intermediate data in each near-memory computing module is less than or equal to a second threshold, that is, the processor evenly distributes the segmented intermediate data to multiple near-memory computing modules for processing.

[0017] In the embodiment of the present application, the processor of the parallel reasoning system can evenly distribute the generated intermediate data to different near-memory computing modules based on the data length of different near-memory computing modules, thereby improving the balance of executing self-attention computing tasks of multiple near-memory computing modules in the parallel reasoning system.

[0018] In one possible implementation, the intermediate data includes a query term matrix and key-value cache data. When the near-memory computing module performs a self-attention calculation task based on the intermediate data, the near-memory computing module performs batch self-attention calculations based on the query term matrix and key-value cache data to generate near-memory calculation results. If the parallel inference system includes multiple near-memory computing modules, the multiple near-memory computing modules can perform self-attention calculations in parallel to generate their own near-memory calculation results. The multiple near-memory computing modules then send their respective near-memory calculation results to the processor.

[0019] The near-memory computing module of the parallel reasoning system in the embodiment of the present application can perform batch self-attention calculations based on the query word matrix and key-value cache data to generate near-memory computing results, thereby improving the feasibility of the solution.

[0020] In one possible implementation, when the parallel reasoning system includes multiple near-memory computing modules, the processor can send the query term matrix in the intermediate data to multiple near-memory computing modules by broadcasting, and the processor can send the key-value cache data in the intermediate data to multiple near-memory computing modules in sequence by appending writes.

[0021] In the embodiment of the present application, the processor of the parallel reasoning system can send different types of intermediate data to the near-memory computing module in different ways, thereby improving the richness of the intermediate data sent by the processor to the near-memory computing module.

[0022] In one possible implementation, the processor and the near-memory computing module can pipeline and parallelize inference tasks for a large language model. The large language model includes multiple transport layers, each of which contains both compute-intensive tasks and self-attention computation tasks. Furthermore, the execution of the compute-intensive task at the current transport layer relies on the near-memory computation results of the previous transport layer. The parallel inference system executes the compute-intensive and self-attention computation tasks in parallel at each transport layer, with the processor executing the compute-intensive task and the near-memory computing module executing the self-attention computation task.

[0023] In the embodiment of the present application, when the parallel reasoning system executes the reasoning task of the large language model, computationally intensive tasks and self-attention calculation tasks can be executed in parallel in the processor and near-memory computing module pipelines respectively, thereby reducing the idle time of the processor and near-memory computing module in the parallel reasoning system and improving the reasoning efficiency of the large language model. The full load operation of the processor and near-memory computing module also reduces the reasoning cost of the large language model.

[0024] In one possible implementation, the processor includes one or more of the following: a graphics processing unit (GPU), a neural network processor (NPU), and a data processing unit (DPU).

[0025] The processors in the parallel reasoning system in the embodiment of the present application can be multiple types of processors, and different types of processors are suitable for different types of reasoning tasks, thereby improving the applicability of the reasoning method.

[0026] In the second aspect, an embodiment of the present application provides an inference device for a large language model, which is applied to a parallel inference system. The parallel inference system includes a processor and a near-memory computing module. Both the processor and the near-memory computing module are used to execute tasks assigned by the parallel inference system. The task processing device includes a transceiver unit and a processing unit. The transceiver unit is used to receive inference tasks of the large language model. The inference tasks include computationally intensive tasks and self-attention computing tasks. The processing unit is used to execute computationally intensive tasks on the processor, generate intermediate data, and send the intermediate data to the near-memory computing module. The computationally intensive tasks include one or more of the following tasks of the large language model: feedforward neural network computing tasks, projection tasks, and layer normalization tasks. The processing unit is also used to execute self-attention computing tasks based on the intermediate data on the near-memory computing module to generate near-memory computing results. The processing unit is also used to generate inference results corresponding to the inference tasks based on the near-memory computing results.

[0027] In one possible implementation, the storage capacity of the near-memory computing module is greater than the storage capacity of the processor, and the computing power of the processor is greater than the computing power of the near-memory computing module.

[0028] In a possible implementation, the length of one or more processing text segments corresponding to the reasoning task is greater than or equal to a first threshold.

[0029] In one possible implementation, the parallel reasoning system includes multiple near-memory computing modules, and the processing unit is used to segment the intermediate data in the processor to obtain multiple segmented intermediate data, and the processor sends the multiple segmented intermediate data to the multiple near-memory computing modules. The processing unit is specifically used to execute self-attention computing tasks in parallel based on the multiple segmented intermediate data in the near-memory computing modules to generate near-memory computing results corresponding to the multiple near-memory computing modules. The processing unit is specifically used to perform forward-reduction computing on the near-memory computing results corresponding to the multiple near-memory computing modules in the processor, and generate an inference result corresponding to the inference task.

[0030] In one possible implementation, the transceiver unit is specifically used to send multiple segmented intermediate data to multiple near-memory computing modules in a processor cycle, so that the difference between the total lengths of the segmented intermediate data in each near-memory computing module is less than or equal to a second threshold.

[0031] In one possible implementation, the intermediate data includes a query term matrix and key-value cache data, and the processing unit is specifically used to perform batch self-attention calculations based on the query term matrix and key-value cache data in the near-memory calculation module to generate near-memory calculation results.

[0032] In one possible implementation, the large language model includes multiple transmission layers, and each of the multiple transmission layers contains computationally intensive tasks and self-attention calculation tasks. Moreover, the execution of computationally intensive tasks in the current transmission layer depends on the near-memory calculation results of the previous transmission layer. The processing unit is also used to execute computationally intensive tasks and self-attention calculation tasks in parallel in each transmission layer. The computationally intensive tasks are executed by the processor, and the self-attention calculation tasks are executed by the near-memory calculation module.

[0033] In one possible implementation, the processor includes one or more of the following: a graphics processing unit (GPU), a neural network processor (NPU), and a data processing unit (DPU).

[0034] In a third aspect, an embodiment of the present application provides a computing device, comprising a processor coupled to a memory, the processor being used to store instructions. When the instructions are executed by the processor, the computing device executes the method described in the first aspect or any possible implementation of the first aspect.

[0035] In a fourth aspect, an embodiment of the present application provides a computing device cluster, which includes one or more computing devices, each of which includes a processor coupled to a memory, and the processor is used to store instructions. When the instructions are executed by the processor, the computing device cluster executes the method described in the first aspect or any possible implementation method of the first aspect.

[0036] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed, the computer executes the method described in the first aspect or any possible implementation method of the first aspect.

[0037] In a sixth aspect, an embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed, the computer implements the method described in the first aspect or any possible implementation method of the first aspect.

[0038] It can be understood that the beneficial effects that can be achieved by any of the above-mentioned inference devices, computing devices, computing device clusters, computer-readable media or computer program products of the large language model can be referred to the beneficial effects in the corresponding methods and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] FIG1 is a schematic diagram of the system architecture of a large language model inference system provided in an embodiment of the present application;

[0040] FIG2 is a flow chart of a large language model inference method provided in an embodiment of the present application;

[0041] FIG3 is a flow chart of another large language model inference method provided in an embodiment of the present application;

[0042] FIG4 is a flow chart of another large language model reasoning method provided in an embodiment of the present application;

[0043] FIG5 is a schematic diagram of a process flow of parallel reasoning of a processor and a near-memory computing module provided in an embodiment of the present application;

[0044] FIG6 is a schematic diagram of a process flow of parallel reasoning of another processor and a near-memory computing module provided in an embodiment of the present application;

[0045] FIG7 is a schematic diagram of the structure of an inference device for a large language model provided in an embodiment of the present application;

[0046] FIG8 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0047] FIG9 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0048] FIG10 is a schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] The embodiments of the present application provide a method and apparatus for reasoning a large language model, which are used to improve the utilization of processors in a parallel reasoning system, thereby improving the reasoning efficiency of the large language model and reducing the reasoning cost of the large language model.

[0050] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0051] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0052] First, some terms involved in the embodiments of the present application are introduced to facilitate those skilled in the art to understand the technical solutions.

[0053] A large language model (LLM) is a natural language processing model based on deep learning technology. Trained using large amounts of text data, it possesses the ability to understand, generate, and process natural language. LLMs treat natural language text as sequential data, such as word or character sequences, and use deep learning models to model the statistical patterns and underlying semantic information in these sequential data. LLMs can perform a wide range of tasks, including text summarization, translation, sentiment analysis, question-answering, and conversation.

[0054] A processor (x-processing unit, XPU) can also be called an accelerator. In this application, a processor refers to a processor such as a graphics processing unit (GPU), a neural processing unit (NPU) and a data processing unit (DPU).

[0055] High-bandwidth memory (HBM) refers to the general term for high-bandwidth memory on the processor XPU, such as the video memory on the graphics processing unit GPU. The neural network processor NPU also has corresponding high-bandwidth storage components. The cost of high-bandwidth memory in the processor is relatively high.

[0056] In order to make the technical solution of the present application clearer and easier to understand, the system architecture of the present application is introduced below with reference to the accompanying drawings.

[0057] Please refer to Figure 1, which is a schematic diagram of the system architecture of a parallel reasoning system provided in an embodiment of the present application. In the system architecture shown in Figure 1, parallel reasoning system 10 is used to perform reasoning tasks for large language models. Parallel reasoning system 10 includes a processor 101 and a near-memory computing module 102. Processor 101 and near-memory computing module 102 are connected via a high-speed transmission bus. Near-memory computing module 102 includes a processor core 1021, memory 1022, and flash memory 1023. The specific functions of each component of large language model reasoning system 10 are described below.

[0058] Processor 101 is used to execute computationally intensive tasks for large language models. These tasks are those that consume a high amount of processor resources during the inference process of large language models, such as feedforward neural network computations, projection tasks, and layer normalization tasks. Processor 101 is also used to send intermediate data generated by executing computationally intensive tasks to the near-memory computing module 102.

[0059] It should be noted that the processor 101 can be a variety of types of processors. For example, the processor 101 can be a graphics processor GPU, a neural network processor NPU or a data processor DPU, without specific limitation.

[0060] The near-memory computing module 102 is used to receive intermediate data sent by the processor 101 and perform self-attention computing tasks based on the intermediate data to generate near-memory computing results. The self-attention computing task is a bandwidth-intensive task, which refers to tasks that consume a large amount of storage resources. Bandwidth-intensive tasks often involve reading and writing large amounts of data, requiring storage space. The near-memory computing module 102 is also used to send the near-memory computing results to the processor 101.

[0061] The near-memory computing module 102 includes a processor core 1021, memory 1022, and flash memory 1023. The processor core 1021 is used to provide computing power for executing self-attention computing tasks. The memory 1022 is used to cache temporary data used by the processor core 1021 during the execution of the self-attention computing tasks. The flash memory 1023 is used to store intermediate data sent by the processor 101 and the near-memory computing results generated by the processor core 1021 during the execution of the self-attention computing tasks.

[0062] It should be noted that the processor 101 and the near-memory computing module 102 in the parallel reasoning system 10 can be deployed in the same computing device or in different computing devices, without any specific limitation.

[0063] In addition, the parallel reasoning system 10 can be a large language model reasoning system composed of multiple processors 101 and multiple near-memory computing modules 102. When the processors 101 and the near-memory computing modules 102 are deployed in different computing devices, the multiple processors 101 or multiple near-memory computing modules 102 in the parallel reasoning system 10 can also be deployed in different computing devices respectively.

[0064] It can be understood that the parallel reasoning system 10 in the embodiment of the present application can be applied to the reasoning scenario of long texts. When the input text of the large language model is long, a large amount of intermediate data will be generated during the reasoning process of the large language model. If these intermediate data are stored in the high-bandwidth memory of the processor 101, the reasoning cost will increase. Therefore, the parallel reasoning system 10 can store this intermediate data based on the near-memory computing module 102 and perform bandwidth-intensive tasks to reduce the reasoning cost of the long text reasoning scenario.

[0065] Based on the parallel reasoning system 10 shown in Figure 1, the present application also provides a large language model reasoning method. The following describes the large language model reasoning method provided in the present application embodiment in conjunction with the embodiment.

[0066] Please refer to Figure 2, which is a flow chart of a method for reasoning a large language model provided in an embodiment of the present application. In the example shown in Figure 2, the method includes the following steps:

[0067] Step 201: The parallel reasoning system receives a reasoning task of a large language model, where the reasoning task includes a computationally intensive task and a self-attention computation task.

[0068] Parallel inference system 10 receives an inference task for a large language model, which includes a computationally intensive task and a self-attention computation task. While executing the inference task for the large language model, parallel inference system 10 assigns the computationally intensive task to processor 101 and the self-attention computation task to near-memory computing module 102.

[0069] It can be understood that since the storage capacity of the near-memory computing module 102 is greater than the storage capacity of the processor 101, the computing power of the processor 101 is greater than the computing power of the near-memory computing module 102. Therefore, the parallel reasoning system 10 will allocate the self-attention computing tasks that occupy high storage resources to the near-memory computing module 102 for execution, thereby improving the utilization rate of the processor 101.

[0070] In one example of step 201, the parallel reasoning system 10 may receive a reasoning task for a large language model. For example, the parallel reasoning system 10 receives a user's input text and infers an output text based on the user's input text. The input text may be a sentence or paragraph of a question asked by the user, and the output text may be an answer to the question, a text summary, or a translation result corresponding to the input text. Specifically, in the process of generating the output text corresponding to the input text, the parallel reasoning system 10 infers the user's input text based on the multi-layer structure of the large language model. Computationally intensive tasks and self-attention computation tasks are required to be performed at each transmission layer of the large language model.

[0071] Please refer to Figure 3, which is a schematic diagram of a large language model reasoning method provided in an embodiment of the present application. In the example shown in Figure 3, the parallel reasoning system 10 receives input text input by the user into the large language model and performs reasoning tasks based on the multi-layer structure of the large language model. In the incremental reasoning (decode) stage of the reasoning task, the parallel reasoning system 10 performs cyclic iterative reasoning in an autoregressive manner at multiple transmission layers of the large language model. In each transmission layer of the large language model, the processor 101 handles computationally intensive tasks, and the near-memory computing module 102 performs self-attention computing tasks.

[0072] It should be noted that in the embodiment of the present application, the large language model can be used to process ultra-long text sequences, that is, the reasoning task of the large language model can correspond to one or more processing texts, and the length of the processing text is greater than or equal to a first threshold, for example, 1,000 characters.

[0073] Since the large language model needs to consume more storage resources when processing ultra-long text sequences, the processor 101 is limited in the length of text sequences that can be processed in batches due to the limitation of high-bandwidth storage, resulting in a relatively low utilization rate of the processor 101. Therefore, the parallel reasoning system 10 offloads the self-attention calculation tasks that occupy a high amount of storage resources to the near-memory computing module 102 for execution, and the processor 101 and the near-memory computing module 102 process the reasoning tasks of the large language model in parallel, thereby improving the utilization rate of the processor 101.

[0074] Step 202: The processor executes computationally intensive tasks, generates intermediate data, and sends the intermediate data to a near-memory computing module.

[0075] After receiving the computationally intensive tasks assigned by the parallel inference system, processor 101 executes the computationally intensive tasks, generates intermediate data, and sends the intermediate data to the near-memory computing module 102. The computationally intensive tasks include one or more of the following tasks for a large language model: feed-forward neural network (FFN) computation, projection, and layer normalization.

[0076] Feedforward network computational tasks include those performed by the feedforward network during the inference process of a large language model, including tasks such as feature extraction and nonlinear transformation. In large language models, feedforward networks typically consist of multiple fully connected layers and activation functions, which are used to perform nonlinear mapping and feature extraction on input data. For example, feedforward networks can perform nonlinear transformations and feature extraction on input word vectors.

[0077] Projection tasks involve mapping low-dimensional word vector representations of input data into a higher-dimensional space. In large language models, projection transforms the model's word vector representations into word vector representations in the output space, enabling further generation of output text. This allows the model to map hidden representations into an appropriate output space, generating results that meet task requirements.

[0078] Layer normalization is the process of normalizing the output of each layer in a large language model. This can accelerate convergence and improve the generalization of large language models. In large language models, layer normalization can be performed as a standard operation for each layer, requiring normalization of the output of each layer.

[0079] In an embodiment of the present application, the intermediate data generated by the processor 101 when performing computationally intensive tasks includes a query term matrix and key-value cache data, wherein the query term matrix is, for example, a matrix composed of query term (query) vectors, and the key-value cache data is, for example, a key (key) vector and a value (value) vector.

[0080] Please continue to refer to Figure 3. In the example shown in Figure 3, during the process of the parallel reasoning system 10 performing the reasoning task of the large language model, in each transmission layer of the large language model, the processor 101 performs computationally intensive tasks such as layer normalization tasks and projection tasks to generate intermediate data, namely the query term matrix Q, key vector K and value vector V.

[0081] In one possible implementation, after the processor 101 performs a computationally intensive task to generate intermediate data, the processor 101 segments the intermediate data to generate a plurality of segmented intermediate data, and sends the plurality of segmented intermediate data to the near-memory computing module 102 for self-attention calculation.

[0082] Please refer to Figure 4, which is a schematic diagram of an inference method for a large language model provided in an embodiment of the present application. In the example shown in Figure 4, the computationally intensive tasks and bandwidth-intensive tasks of the large language model are executed in processor 101 and near-memory computing module 102, respectively. In each round of inference calculation of the large language model, processor 101 needs to send intermediate data to near-memory computing module 102 and receive the near-memory calculation results sent by near-memory computing module 102.

[0083] For example, in the example shown in Figure 4, the processor 101 performs projection tasks and feedforward network calculation tasks based on computationally intensive operators during the inference process of each transmission layer, generates intermediate data for reasoning at each layer, and sends the intermediate data to the near-memory calculation module 102 during the inference process of each layer. The near-memory calculation module 102 performs self-attention calculation tasks based on bandwidth-intensive operators, generates near-memory calculation results for reasoning at each layer, and sends the near-memory calculation results to the processor 101.

[0084] In the example shown in FIG4 , the processor 101 also segments the generated intermediate data and sends the segmented intermediate data to the near-memory computing module. For example, the processor 101 segments the key-value cache generated by executing computationally intensive tasks such as layer normalization and projection tasks into key-value cache data KV1, key-value cache data KV2, and key-value cache data KV2, and sends the key-value cache data KV1, key-value cache data KV2, and key-value cache data KV2 to the near-memory computing module 102 for self-attention calculation.

[0085] In one possible implementation, if the parallel reasoning system 10 includes multiple near-memory computing modules 102, in the process of the processor 101 sending segmented intermediate data to the multiple near-memory computing modules 102, the processor 101 cyclically sends multiple segmented intermediate data to the multiple near-memory computing modules 102 respectively, wherein the difference between the total lengths of the segmented intermediate data in the multiple near-memory computing modules 102 is less than or equal to the second threshold, that is, the segmented intermediate data of the processor 101 is evenly distributed among the multiple near-memory computing modules 102 for processing.

[0086] It should be noted that if the parallel reasoning system 10 includes multiple near-memory computing modules 102, then for the query word matrix in the intermediate data, the processor 101 can send it to multiple near-memory computing modules 102 by broadcasting, and for the key-value cache data in the intermediate data, the processor 101 can send it to multiple near-memory computing modules 102 in sequence by appending.

[0087] Please continue to refer to Figure 3. In the example shown in Figure 3, the parallel reasoning system 10 includes two near-memory computing modules, namely near-memory computing module 1 and near-memory computing module 2. The processor 101 performs computationally intensive tasks to generate a query term matrix Q and key-value cache data. The key-value cache data includes a key vector K and a value vector V. The processor 101 sends the query term matrix Q and the key-value cache data to the near-memory computing module 1 and the near-memory computing module 2 respectively.

[0088] In the example shown in FIG3 , during the prefill inference phase of the large language model, processor 101 has already stored some key-value cache data segments in near-memory computing module 1 and near-memory computing module 2. For example, near-memory computing module 1 stores key-value segment KV segment0, and near-memory computing module 2 stores key-value segment KV segment1. During the incremental inference (decode) phase of the large language model, processor 101 sends the newly added key-value cache data to near-memory computing module 1 and near-memory computing module 2 in an append-write manner. To prevent an imbalance in the amount of computation between the two near-memory computing modules, processor 101 segments the key-value cache data and cyclically sends the key-value cache data segments to near-memory computing module 1 and near-memory computing module 2. This cyclical sending method is also called round-robin allocation.

[0089] For example, in the example shown in Figure 3, when the processor 101 sends KV segment 3 to the near memory computing module 2, if the total length of the key-value cache data in the near memory computing module 2 exceeds the preset value, the processor 101 stops sending the key-value cache data to the computing module 2 and re-sends KV segment 4 to the near memory computing module 1, so that the difference in the total length of the key-value cache data segments between the near memory computing module 1 and the near memory computing module 2 is less than the second threshold.

[0090] In the example shown in FIG3 , for the query word matrix Q generated by the processor 101 when executing the computationally intensive task, the processor 101 sends the query word matrix Q to the near-storage computing module 1 and the near-storage computing module 2 in a broadcast manner.

[0091] Step 203. The near-memory calculation module performs the self-attention calculation task based on the intermediate data and generates the near-memory calculation results.

[0092] After receiving the intermediate data sent by the processor 101, the near-memory calculation module 102 performs the self-attention calculation task based on the intermediate data, generates the near-memory calculation results, and sends the near-memory calculation results to the processor 101. Specifically, the intermediate data includes the query word matrix Q and the key-value cache data KV. The near-memory calculation module 102 performs batch self-attention calculations based on the query word matrix Q and the key-value cache data KV to generate the near-memory calculation results.

[0093] It is understood that if the parallel reasoning system 10 includes multiple near-memory computing modules 102 , the multiple near-memory computing modules 102 can perform self-attention computations in parallel to generate their own near-memory computation results. The multiple near-memory computing modules 102 send their respective near-memory computation results to the processor 101 .

[0094] Continuing with FIG3 , in the example shown in FIG3 , after near-memory computing module 1 and near-memory computing module 2 respectively receive the query term matrix Q and key-value cache data sent by processor 101 , near-memory computing module 1 and near-memory computing module 2 perform batched self-attention calculations based on their respective KV segments to generate near-memory calculation results, such as score1 for near-memory computing module 1 and score2 for near-memory computing module 2. Near-memory computing module 1 and near-memory computing module 2 each send their respective near-memory calculation results to processor 101 .

[0095] Step 204: The processor generates an inference result corresponding to the inference task based on the nearby calculation result.

[0096] After receiving the near-memory computation results sent by the near-memory computation module 102, the processor 101 generates an inference result corresponding to the inference task based on the near-memory computation results. If the parallel inference system 10 includes multiple near-memory computation modules 102, the processor 101 performs a forward-reduction computation on the near-memory computation results corresponding to the multiple near-memory computation modules 102 and generates an inference result corresponding to the inference task.

[0097] Please refer to 3. In the example shown in FIG3, after processor 101 receives the near-memory calculation results score1 and score2 sent by near-memory calculation module 1 and near-memory calculation module 2, processor 101 performs a reduction calculation on the near-memory calculation results score1 and score2 and sums the residuals to obtain a processed near-memory calculation result. Processor 101 then proceeds to execute the computationally intensive task of the next transport layer based on the processed near-memory calculation result.

[0098] In one possible implementation, the processor 101 and the near-memory computing module 102 can pipeline and process inference tasks for a large language model. Specifically, because the large language model includes multiple transmission layers, each of the multiple transmission layers contains computationally intensive tasks and self-attention computation tasks, and the execution of computationally intensive tasks at the current transmission layer depends on the near-memory computation results of the previous transmission layer, the parallel inference system 10 executes computationally intensive tasks and self-attention computation tasks in parallel at each transmission layer, wherein the computationally intensive tasks are executed by the processor 101 and the self-attention computation tasks are executed by the near-memory computing module 102.

[0099] Please refer to Figure 5, which is a schematic diagram of another large language model reasoning method provided in an embodiment of the present application. In the example shown in Figure 5, the processor 101 and the near-memory computing module 102 in the parallel reasoning system 10 respectively execute the feedforward network computing task and the self-attention computing task in parallel, wherein the processor 101 executes the feedforward network computing task at the current transmission layer, which requires the attention computing result of the near-memory computing module 102 at the previous transmission layer. For example, in Figure 5, the processor 101 executing the FFN task needs to rely on the near-memory computing result of the Attn task executed by the near-memory computing module 102 in the previous layer.

[0100] In the example shown in Figure 5, in order to reduce the idle time of the processor 101 and the near memory computing module 102, the processor 101 and the near memory computing module 102 respectively use a pipeline parallel manner to execute the feedforward network computing task and the self-attention computing task, that is, when the processor 101 executes the feedforward network computing task corresponding to the previous transmission layer, the near memory computing module 102 executes the attention computing task corresponding to the next transmission layer in parallel.

[0101] For example, in the example shown in Figure 5, when the processor 101 is executing the computationally intensive task of the reasoning task req1, the near memory computing module 102 is executing the self-attention computing task of the reasoning task req2 in parallel. When the processor 101 is executing the computationally intensive task of the reasoning task req2, the near memory computing module 102 is executing the self-attention computing task of the next transmission layer of the reasoning task req1 in parallel, thereby avoiding the processor 101 and the near memory computing module 102 from being idle.

[0102] Please refer to Figure 6, which is a schematic diagram of pipelined parallel reasoning for a large language model provided in an embodiment of the present application. In the example shown in Figure 6, the processor 101 and the near-memory computing module 102 respectively execute the feedforward network computing task and the self-attention computing task in parallel, wherein the reasoning task rep1 and the reasoning task req3 are processed alternately by the processor 101 and the near-memory computing module 102. While the near-memory computing module 102 executes the self-attention computing task corresponding to rep1, the processor 101 executes the feedforward network computing task corresponding to rep3 in parallel.

[0103] For example, in the example shown in Figure 6, between time S+0 and time S+1, the near memory computing module 102 is executing the self-attention computing tasks Attn1-1 and Attn1-2 and the self-attention computing tasks Attn2-1 and Attn2-2 while the processor 101 is executing the feedforward network computing tasks FNN3-1 and FNN4-1 in parallel.

[0104] In the example shown in FIG6 , the processor 101 may also execute multiple reasoning tasks in parallel. For example, the processor 101 executes reasoning tasks rep3 and rep4 in parallel between time S+0 and time S+1. The processor 101 executes reasoning tasks rep1 and rep2 in parallel between time S+1 and time S+2.

[0105] It can be seen from the above embodiments that the processor of the parallel reasoning system in the embodiments of the present application can generate intermediate data for executing computationally intensive tasks and send it to the near-memory computing module, so that the near-memory computing module executes wide-intensive self-attention computing tasks based on the intermediate data, thereby offloading the self-attention computing tasks to the near-memory computing module for execution, so that the processor of the parallel reasoning system can avoid memory limitations to execute computationally intensive tasks, thereby improving the utilization of the processor.

[0106] Based on the above method embodiment, the embodiment of the present application also provides an inference device for a large language model. The following specifically introduces the inference device for a large language model provided by the embodiment of the present application.

[0107] Please refer to Figure 7, which is a schematic diagram of the structure of an inference device for a large language model provided in an embodiment of the present application. In the example shown in Figure 7, the inference device 700 for a large language model is used to implement the various steps performed by the parallel inference system in the above-mentioned embodiments. The inference device 700 for a large language model includes a transceiver unit 701 and a processing unit 702.

[0108] Among them, the transceiver unit 701 is used to receive the inference tasks of the large language model, and the inference tasks include computationally intensive tasks and self-attention calculation tasks. The processing unit 702 is used to execute computationally intensive tasks on the processor, generate intermediate data, and send the intermediate data to the near-memory computing module. The computationally intensive tasks include one or more of the following tasks of the large language model: feedforward neural network computing tasks, projection tasks, and layer normalization tasks. The processing unit 702 is also used to execute the self-attention computing tasks based on the intermediate data in the near-memory computing module to generate near-memory computing results. The processing unit 702 is also used to generate inference results corresponding to the inference tasks based on the near-memory computing results.

[0109] In one possible implementation, the storage capacity of the near-memory computing module is greater than the storage capacity of the processor, and the computing power of the processor is greater than the computing power of the near-memory computing module.

[0110] In a possible implementation, the length of one or more processing text segments corresponding to the reasoning task is greater than or equal to a first threshold.

[0111] In one possible implementation, the parallel reasoning system includes multiple near-memory computing modules, and the processing unit 702 is used to segment the intermediate data in the processor to obtain multiple segmented intermediate data, and the processor sends the multiple segmented intermediate data to the multiple near-memory computing modules. The processing unit 702 is specifically used to execute the self-attention calculation task in parallel based on the multiple segmented intermediate data in the near-memory computing module, and generate the near-memory calculation results corresponding to the multiple near-memory computing modules. The processing unit 702 is specifically used to perform forward-reduction calculations on the near-memory calculation results corresponding to the multiple near-memory computing modules in the processor, and generate the reasoning result corresponding to the reasoning task.

[0112] In one possible implementation, the transceiver unit 701 is specifically used to send multiple segmented intermediate data to multiple near-memory computing modules in a cyclic manner in the processor, so that the difference between the total lengths of the segmented intermediate data in each near-memory computing module is less than or equal to a second threshold.

[0113] In one possible implementation, the intermediate data includes a query term matrix and key-value cache data, and the processing unit 702 is specifically used to perform batch self-attention calculations based on the query term matrix and key-value cache data in the near-memory calculation module to generate near-memory calculation results.

[0114] In one possible implementation, the large language model includes multiple transmission layers, and each of the multiple transmission layers has computationally intensive tasks and self-attention calculation tasks. In addition, the execution of computationally intensive tasks in the current transmission layer depends on the near-memory calculation results of the previous transmission layer. The processing unit 702 is also used to execute computationally intensive tasks and self-attention calculation tasks in parallel in each transmission layer. The computationally intensive tasks are executed by the processor, and the self-attention calculation tasks are executed by the near-memory calculation module.

[0115] In one possible implementation, the processor includes one or more of the following: a graphics processing unit (GPU), a neural network processor (NPU), and a data processing unit (DPU).

[0116] It can be understood that the transceiver unit 701 and the processing unit 702 in the large language model reasoning device 700 can be mapped as functional modules to the various modules in the large language model reasoning system 10 in Figure 1, thereby realizing the functions of the various modules in the large language model reasoning system 10.

[0117] It should be understood that the division of units in the above device is merely a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. Moreover, the units in the device can all be implemented in the form of software called through processing elements; or they can all be implemented in the form of hardware; or some units can be implemented in the form of software called through processing elements, and some units can be implemented in the form of hardware. For example, each unit can be a separately established processing element, or it can be integrated into a certain chip of the device. In addition, it can also be stored in the memory in the form of a program, called by a certain processing element of the device and perform the function of the unit. In addition, all or part of these units can be integrated together, or they can be implemented independently. The processing element described here can also be a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each unit above can be implemented by the integrated logic circuit of the hardware in the processor element or in the form of software called through the processing element.

[0118] It is worth noting that, for the sake of simplicity of description, the above method embodiments are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited to the order of the actions described. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required for this application.

[0119] Other reasonable step combinations that can be thought of by those skilled in the art based on the above description also fall within the scope of protection of this application. Secondly, those skilled in the art should also be familiar with that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.

[0120] Please refer to Figure 8, which is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. As shown in Figure 8, the computing device 800 includes: a processor 801, a memory 802, a communication interface 803, and a bus 804. The processor 801, the memory 802, and the communication interface 803 are coupled via a bus (not labeled in the figure). The memory 802 stores instructions. When the execution instructions in the memory 802 are executed, the computing device 800 performs the method performed by the computing device in the above method embodiment.

[0121] The computing device 800 may be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASICs), one or more digital signal processors (DSPs), one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms. For example, when a unit in the device can be implemented in the form of a processing element scheduler, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call a program. For example, these units may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0122] The processor 801 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0123] Memory 802 may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0124] Memory 802 stores executable program code, which processor 801 executes to implement the functions of the aforementioned units or modules, thereby implementing the aforementioned large language model inference method. That is, memory 802 stores instructions for executing the aforementioned large language model inference method.

[0125] The communication interface 803 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 800 and other devices or a communication network.

[0126] In addition to the data bus, bus 804 may also include a power bus, a control bus, and a status signal bus. The bus may be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a unified bus (Ubus or UB), a Compute Express Link (CXL), or a Cache Coherent Interconnect for Accelerators (CCIX). Buses can be categorized as address buses, data buses, and control buses.

[0127] Please refer to FIG9 , which is a schematic diagram of a computing device cluster provided in an embodiment of the present application. As shown in FIG9 , the computing device cluster 900 includes at least one computing device 800 .

[0128] 9 , the computing device cluster 900 includes at least one computing device 800. The memory 802 in one or more computing devices 800 in the computing device cluster 900 may store the same instructions for executing the above-mentioned inference method for the large language model.

[0129] In some possible implementations, the memory 802 of one or more computing devices 800 in the computing device cluster 900 may also store partial instructions for executing the aforementioned large language model reasoning method. In other words, the combination of one or more computing devices 800 can collectively execute instructions for executing the aforementioned large language model reasoning method.

[0130] It should be noted that the memory 802 in different computing devices 800 in computing device cluster 900 can store different instructions, each for executing a portion of the functions of the aforementioned large language model inference apparatus. In other words, the instructions stored in the memory 802 in different computing devices 800 can implement the functions of one or more modules in the transceiver unit and the processing unit.

[0131] In some possible implementations, one or more computing devices 800 in the computing device cluster 900 may be connected via a network, which may be a wide area network or a local area network.

[0132] Please refer to Figure 10, which is a schematic diagram of computer devices in a computer cluster provided by an embodiment of the present application connected via a network. As shown in Figure 10, two computing devices 800A and 800B are connected via a network. Specifically, the connection to the network is through a communication interface in each computing device.

[0133] In one possible implementation, the memory of the computing device 800A stores instructions for executing the functions of the transceiver unit, while the memory of the computing device 800B stores instructions for executing the functions of the processing unit.

[0134] It should be understood that the functions of the computing device 800A shown in Figure 10 may also be completed by multiple computing devices. Similarly, the functions of the computing device 800B may also be completed by multiple computing devices.

[0135] In another embodiment of the present application, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When the processor of the device executes the computer-executable instructions, the device executes the method executed by the computing device in the above method embodiment.

[0136] In another embodiment of the present application, a computer program product is provided, the computer program product including computer-executable instructions stored in a computer-readable storage medium. When a processor of a device executes the computer-executable instructions, the device performs the method performed by the computing device in the above method embodiment.

[0137] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0138] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0139] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0140] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0141] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. An inference method for a large language model, characterized in that, Applied to a parallel inference system, the parallel inference system includes a processor and a near-memory computing module, both the processor and the near-memory computing module are used to execute tasks assigned by the parallel inference system, and the method includes: The parallel inference system receives an inference task of a large language model, and the inference task includes a computationally intensive task and a self-attention calculation task; The processor executes the computationally intensive task, generates intermediate data, and sends the intermediate data to the near-memory computing module. The computationally intensive task includes one or more of the following tasks of the large language model: feed-forward neural network calculation task, projection task, and layer normalization task; The near-memory computing module executes the self-attention calculation task based on the intermediate data to generate a near-memory calculation result; The processor generates an inference result corresponding to the inference task based on the near-memory calculation result.

2. The method according to claim 1, wherein The storage capacity of the near-memory computing module is greater than the storage capacity of the processor, and the computing power of the processor is greater than the computing power of the near-memory computing module.

3. The method according to claim 1 or 2, characterized in that, One or more segments of processed text corresponding to the inference task, and the length of the processed text is greater than or equal to a first threshold.

4. The method according to any one of claims 1 to 3, characterized in that The parallel inference system includes multiple near-memory computing modules, and the processor sending the intermediate data to the near-memory computing module includes: The processor segments the intermediate data to obtain multiple segmented intermediate data; The processor sends the multiple segmented intermediate data to the multiple near-memory computing modules; The near-memory computing module executing the self-attention calculation task based on the intermediate data includes: The multiple near-memory computing modules execute the self-attention calculation task in parallel based on the multiple segmented intermediate data to generate near-memory calculation results corresponding to the multiple near-memory computing modules; The processor generating an inference result corresponding to the inference task based on the near-memory calculation result includes: The processor performs a reduction calculation on the near-memory calculation results corresponding to the multiple near-memory computing modules and generates an inference result corresponding to the inference task.

5. The method according to claim 4, wherein The processor sending the multiple segmented intermediate data to the multiple near-memory computing modules includes: The processor circulates and sends the multiple segmented intermediate data to the multiple near-memory computing modules respectively, so that the difference between the total lengths of the segmented intermediate data in each near-memory computing module is less than or equal to a second threshold.

6. The method according to any one of claims 1 to 3, characterized in that, The intermediate data includes a query word matrix and key-value cache data, and the near-memory computing module executing the self-attention calculation task based on the intermediate data further includes: The near-memory computing module performs batch self-attention calculation based on the query word matrix and the key-value cache data to generate the near-memory calculation result.

7. The method according to any one of claims 1 to 4, characterized in that, The large language model includes multiple transmission layers, and there are both the computationally intensive task and the self-attention calculation task in each of the multiple transmission layers. Moreover, executing the computationally intensive task in the current transmission layer depends on the near-memory calculation result of the previous transmission layer. The method further includes: The parallel inference system executes the computationally intensive task and the self-attention calculation task in parallel at each transport layer. The computationally intensive task is executed by the processor, and the self-attention calculation task is executed by the near-memory computing module.

8. The method according to any one of claims 1 to 5, characterized in that The processor includes one or more of the following: a graphics processing unit (GPU), a neural network processing unit (NPU), and a data processing unit (DPU).

9. An inference device for a large language model, characterized in that, Applied to a parallel inference system, the parallel inference system includes a processor and a near-memory computing module. Both the processor and the near-memory computing module are used to execute the tasks assigned to the parallel inference system. The apparatus includes: A transceiver unit for receiving an inference task of a large language model, where the inference task includes a computationally intensive task and a self-attention calculation task; A processing unit for executing the computationally intensive task in the processor, generating intermediate data, and sending the intermediate data to the near-memory computing module. The computationally intensive task includes one or more of the following tasks of the large language model: a feed-forward neural network calculation task, a projection task, and a layer normalization task; The processing unit is further configured to execute the self-attention calculation task in the near-memory computing module based on the intermediate data, generating a near-memory computing result; The processing unit is further configured to generate an inference result corresponding to the inference task based on the near-memory computing result.

10. The device according to claim 9, characterized in that, The storage capacity of the near-memory computing module is greater than the storage capacity of the processor, and the computing power of the processor is greater than the computing power of the near-memory computing module.

11. The device according to claim 9 or 10, characterized in that, One or more segments of processed text corresponding to the inference task, where the length of the processed text is greater than or equal to a first threshold.

12. The device according to any one of claims 9 to 11, characterized in that The parallel inference system includes multiple near-memory computing modules. The processing unit is configured to: Segment the intermediate data in the processor to obtain multiple segmented intermediate data; Send the multiple segmented intermediate data to the multiple near-memory computing modules in the processor; Specifically, the processing unit is configured to: Execute the self-attention calculation task in parallel in the near-memory computing module based on the multiple segmented intermediate data, generating near-memory computing results corresponding to the multiple near-memory computing modules; Specifically, the processing unit is configured to: Perform a reduction calculation on the near-memory computing results corresponding to the multiple near-memory computing modules in the processor and generate an inference result corresponding to the inference task.

13. The device according to claim 12, characterized in that, Specifically, the transceiver unit is configured to: Circulate and send the multiple segmented intermediate data to the multiple near-memory computing modules in the processor respectively, so that the difference between the total lengths of the segmented intermediate data in each near-memory computing module is less than or equal to a second threshold.

14. The device according to any one of claims 9 to 11, characterized in that The intermediate data includes a query word matrix and key-value cache data. Specifically, the processing unit is configured to: Perform a batch self-attention calculation in the near-memory computing module based on the query word matrix and the key-value cache data, generating the near-memory computing result.

15. The device according to any one of claims 9 to 12, characterized in that, The large language model includes multiple transmission layers, and both the computationally intensive tasks and the self-attention calculation tasks exist in each of the multiple transmission layers. Moreover, the execution of the computationally intensive tasks in the current transmission layer depends on the near-memory calculation results of the previous transmission layer. The processing unit is further configured to: Execute the computationally intensive tasks and the self-attention calculation tasks in parallel in each of the transmission layers. The computationally intensive tasks are executed by the processor, and the self-attention calculation tasks are executed by the near-memory calculation module.

16. The device according to any one of claims 9 to 15, characterized in that The processor includes one or more of the following: a graphics processing unit (GPU), a neural network processing unit (NPU), and a data processing unit (DPU).

17. A computing device, characterized in that, It includes a processor, the processor is coupled to a memory, and the processor is configured to store instructions. When the instructions are executed by the processor, the electronic device is caused to execute the method according to any one of claims 1 to 8.

18. A cluster of computing devices, characterized in that, It includes at least one computing device, the computing device includes a processor, the processor is coupled to a memory, and the processor is configured to store instructions. When the instructions are executed by the processor, the computing device cluster is caused to execute the method according to any one of claims 1 to 8.

19. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instructions are executed, the computer is caused to execute the method according to any one of claims 1 to 8.

20. A computer program product comprising instructions, characterized in that, When the instructions are executed, the computer is caused to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Big language model reasoning method and device

    CN120338089A

  • Near memory calculation device, near memory calculation method, integrated circuit, and storage medium

    CN116089356A

  • Large language model reasoning system, method and equipment without perception of server

    CN116702907A

  • Text reasoning task processing method and device, equipment and storage medium

    CN116822629A

  • Large language model reasoning optimization method and device, computer equipment and storage medium

    CN117194056A

Cited By

  • Data processing system, method, device, medium and program product

    CN120743554A

  • Request inference task processing method, device, equipment, medium, product and heterogeneous system

    CN121233344A

  • Method for reasoning by using large language model, control device and storage medium

    CN121303273A