FPGA-based high-performance large language model accelerator and reasoning method

By designing a multi-computing unit and matrix processing unit accelerator on the FPGA platform, combining parallel computing and memory scheduling mechanisms, the problems of low computing efficiency, insufficient memory bandwidth, high latency and low energy efficiency in the existing technology are solved, and an efficient, low latency and energy-saving large-language model inference solution is realized.

CN119990213APending Publication Date: 2025-05-13BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510157119.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing FPGA acceleration solutions fail to fully utilize the advantages of parallel computing power and memory management when processing large models, resulting in insufficient computing efficiency and memory bandwidth, low latency and energy efficiency.

Method used

A high-performance large language model accelerator based on FPGA is designed, and a structure that combines multiple computing units and matrix processing units is connected to the host CPU through a PCIe bus and connected to high-bandwidth memory. The parallel computing and memory scheduling mechanism is adopted to optimize data loading, computing and storage procedures, dynamically allocate memory resources, optimize data storage location, and ensure optimal data access for each computing unit under memory bandwidth limitation.

Benefits of technology

It significantly improves computing efficiency, improves memory bandwidth utilization, reduces latency and energy consumption, and realizes a high-efficiency, low-latency and energy-saving large-language model inference solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990213A_ABST
    Figure CN119990213A_ABST
Patent Text Reader

Abstract

The invention discloses an FPGA (Field Programmable Gate Array)-based high-performance large language model accelerator and a reasoning method, which adopt a structure of combining a plurality of computing units (CU) and a matrix processing unit (MPE), and can efficiently allocate computing tasks and realize parallel processing by virtue of the parallel computing capability of the FPGA, thereby greatly improving the computing efficiency. Through parallel processing of a plurality of tasks, the reasoning speed is obviously improved, and the calculation bottleneck in a traditional scheme is avoided. And secondly, in the aspect of memory management, through a hybrid storage strategy of a high-bandwidth memory (HBM) and an off-chip memory (such as DDR), utilization of memory bandwidth is optimized, efficient flowing of data among a plurality of computing units is realized, delay in data transmission is reduced, and efficient data access in the computing process is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of high-performance large language model accelerators based on FPGA, and in particular to a high-performance large language model accelerator and reasoning method based on FPGA. Background Art

[0002] At present, computing methods based on GPUs (graphics processing units) and traditional hardware can no longer meet the application requirements of large models in high-performance, low-latency, and low-energy consumption environments. Although GPUs have powerful parallel computing capabilities, they still face computing efficiency and memory bandwidth bottlenecks when processing large-scale AI models. In addition, GPUs are not suitable for certain types of computing tasks, such as sparse computing and mixed precision quantization of large models, which usually have high requirements for hardware customization and optimization. Therefore, when facing large-scale neural network inference tasks, the existing hardware architecture is difficult to balance multiple requirements such as high efficiency, low latency, and low energy consumption, especially in resource-constrained environments such as edge computing, the Internet of Things (IoT), and satellite communications.

[0003] In order to make up for the shortcomings of traditional GPUs in large-scale AI model acceleration, FPGA (field programmable gate array), as a programmable hardware platform, has become an effective acceleration solution with its high customizability and parallel computing capabilities. FPGA can perform hardware-level optimization according to specific computing requirements, and is particularly suitable for sparse computing and quantization operations in deep learning, and has higher energy efficiency than GPU. However, although FPGA has shown good performance in the field of hardware acceleration, existing FPGA acceleration solutions usually have the following problems: First, traditional FPGA accelerators fail to fully utilize the parallel computing capabilities and memory management advantages of FPGA when processing large models, resulting in a large room for improvement in computing efficiency and throughput; second, due to the limitations of FPGA hardware resources, the performance of existing solutions in memory bandwidth and data flow optimization is still not ideal, and cannot effectively reduce the delay and energy consumption in the computing process; finally, the existing FPGA acceleration solution has not yet provided a comprehensive solution for the efficiency and scalability of large-scale model reasoning.

[0004] Therefore, how to achieve efficient large-model inference acceleration on the FPGA platform and solve problems such as computing efficiency, memory bandwidth, latency, and energy efficiency in existing technologies has become a key challenge that needs to be urgently addressed in the current technological development. Summary of the invention

[0005] The present application provides a high-performance large language model accelerator and reasoning method based on FPGA, aiming to solve the problems of computing efficiency, memory bandwidth, latency and energy efficiency existing in the prior art.

[0006] In a first aspect, a high-performance large language model accelerator based on FPGA is provided, the accelerator comprising:

[0007] A plurality of computing units, wherein the computing units are connected to a host CPU via a PCIe bus and are connected to a high bandwidth memory;

[0008] Each of the computing units includes at least one matrix processing unit and performs parallel computing according to a periodic process of loading-computing-storing;

[0009] The accelerator optimizes the data loading, computing and storage processes through parallel computing among multiple computing units and their memory scheduling mechanism, wherein the multiple computing units work in coordination to ensure the efficient operation of each computing unit when loading data, executing computing tasks and storing results; the memory scheduling mechanism is responsible for dynamically allocating memory resources, optimizing data storage locations and ensuring optimal data access for each computing unit under memory bandwidth constraints.

[0010] In the above scheme, optionally, the matrix processing unit includes multiple configurable DSP cores, which are configured through HLS technology to achieve efficient parallel processing of matrix operations required in a large language model.

[0011] In the above solution, optionally, each computing unit in the accelerator communicates with the high-bandwidth memory through an independent AXI interface, and each computing unit can flexibly schedule the memory bandwidth to optimize the utilization of computing resources.

[0012] In the above solution, optionally, the memory management module of the accelerator includes multiple read buffers and write buffers, supports multi-channel parallel access to memory, and realizes parallel transmission of data through DMA channels to reduce data transmission delay.

[0013] In the above scheme, optionally, the accelerator reduces the idle time in the load-compute-store cycle through an optimized data flow structure, ensuring that each computing unit runs efficiently, thereby maximizing the computing throughput.

[0014] In the above solution, optionally, the memory management module of the accelerator dynamically allocates data storage between on-chip memory and off-chip memory according to data flow and computing requirements.

[0015] In the above solution, optionally, the accelerator optimizes the data transmission path to ensure that data is quickly and efficiently transmitted from the host CPU to the FPGA, thereby reducing the overall delay of the system.

[0016] In the above scheme, optionally, the accelerator adopts an FPGA platform and is implemented on Xilinx Alveo U280 hardware, optimizing the reasoning process of the large language model through its programmable hardware resources.

[0017] In the above scheme, optionally, the accelerator can adjust the configuration of the computing unit according to the characteristics of the large language model to adapt to different types of reasoning tasks and achieve hardware-level acceleration optimization.

[0018] In a second aspect, a method for accelerating large language model reasoning based on FPGA is characterized in that the method comprises the following steps:

[0019] Generate large language model inference tasks in the host CPU and transfer the data to the FPGA through the PCIe interface;

[0020] Using multiple computing units in the FPGA to perform inference tasks in parallel, each computing unit includes a matrix processing unit;

[0021] Use high-bandwidth memory to provide data to the computing unit and read and write data efficiently through parallel memory access;

[0022] In the load-compute-store cycle, the computing throughput of each computing unit is optimized and the idle time is reduced.

[0023] Compared with the prior art, this application has at least the following beneficial effects:

[0024] Based on further analysis and research of the problems of the prior art, this application recognizes that the prior art has problems such as computing efficiency, memory bandwidth, delay and energy efficiency. Through unique hardware architecture design and efficient computing resource management, it significantly solves the problems of low computing efficiency, memory bandwidth bottleneck, high delay and low energy efficiency mentioned in the background technology. First, by adopting a structure combining multiple computing units (CU) and matrix processing units (MPE), with the help of the parallel computing capability of FPGA, it is possible to efficiently allocate computing tasks and realize parallel processing, thereby greatly improving the computing efficiency. In traditional hardware platforms such as GPU, due to the huge computing tasks and the problem of waste of computing resources, the present invention significantly improves the reasoning speed by processing multiple tasks in parallel, avoiding the computing bottleneck in the traditional solution. Secondly, in terms of memory management, the present invention optimizes the utilization of memory bandwidth through a hybrid storage strategy of high bandwidth memory (HBM) and off-chip memory (such as DDR), solving the problem of insufficient memory bandwidth often occurring in traditional hardware when processing large-scale data. Through multiple buffers and DMA data transmission mechanism in the memory management module, efficient flow of data between multiple computing units is achieved, delays in data transmission are reduced, and efficient data access during the calculation process is ensured. Thirdly, the present invention adopts a load-compute-store parallel cycle strategy to avoid idle time of computing resources, further improve computing throughput and utilization efficiency of hardware resources, and significantly reduce latency problems in traditional hardware. In addition, the programmability and hardware customization capabilities of the FPGA platform enable the present invention to dynamically adjust the use of computing resources and memory according to different computing requirements, thereby improving overall energy efficiency. Compared with traditional hardware platforms such as GPUs, it can reduce power consumption while maintaining high computing performance.

[0025] In summary, the present invention solves many difficult problems raised in the background technology by optimizing the design of computing units, memory management, data flow and computing cycles, and provides an efficient, low-latency and energy-saving large language model reasoning solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 An architectural block diagram of a high-performance large language model accelerator based on FPGA provided in one embodiment of the present application;

[0027] Figure 2 A Tinyllama structure diagram provided for one embodiment of the present application;

[0028] Figure 3 A hardware architecture block diagram of a parallel matrix processing unit provided in one embodiment of the present application;

[0029] Figure 4 A block diagram of a hybrid storage architecture provided for one embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0031] In one embodiment, Figure 1 As shown, a high-performance large language model accelerator based on FPGA is provided, and the accelerator includes:

[0032] A plurality of computing units, wherein the computing units are connected to a host CPU via a PCIe bus and are connected to a high bandwidth memory;

[0033] Each of the computing units includes at least one matrix processing unit and performs parallel computing according to a periodic process of loading-computing-storing;

[0034] The accelerator optimizes the data loading, computing and storage processes through parallel computing among multiple computing units and their memory scheduling mechanism, wherein the multiple computing units work in coordination to ensure the efficient operation of each computing unit when loading data, executing computing tasks and storing results; the memory scheduling mechanism is responsible for dynamically allocating memory resources, optimizing data storage locations and ensuring optimal data access for each computing unit under memory bandwidth constraints.

[0035] The FPGA accelerator in this embodiment adopts a multi-computing unit (CU) structure, which is connected to the host CPU through the PCIe bus and to the high-bandwidth memory (HBM). Each computing unit is responsible for parallel processing of various reasoning tasks in the large language model to optimize computing efficiency. Each computing unit is equipped with at least one matrix processing unit (MPE), and MPE implements efficient matrix operations through multiple configurable DSP cores.

[0036] The parallelism of these computing units can significantly improve the processing power of reasoning tasks, reduce processing time, and thus reduce latency. Multiple computing units work together to simultaneously process multiple computing tasks in a large language model, which is especially important for AI reasoning tasks that require large-scale parallel computing.

[0037] The matrix processing unit (MPE) is a key computing module in the accelerator of this embodiment, which is responsible for efficiently processing matrix operations of large language models. Each MPE includes multiple DSP cores, which are optimized by high-level synthesis (HLS) technology to support efficient parallel computing. Matrix operations usually involve a large number of data multiplication and addition operations. By configuring multiple parallel DSP blocks in the MPE, the present invention can greatly improve the throughput and efficiency of the calculation.

[0038] Each MPE processes multiple computing tasks in parallel, reducing the waiting time between tasks, thereby maximizing the resource utilization of the hardware. Through parallel computing, the DSP core in the computing unit can process multiple computing tasks at the same time, reducing the computing bottleneck existing in traditional single-core computing platforms.

[0039] In order to optimize the use of memory bandwidth, this embodiment adopts a hybrid storage strategy of high bandwidth memory (HBM) and off-chip memory (such as DDR). The memory management module of the accelerator includes multiple buffers (such as buffer1 to buffer6), and parallel data transmission between memories is realized through DMA channels. This memory management solution ensures efficient use of memory and reduces computing pauses caused by memory access delays.

[0040] When data flows between different computing units and memory, an efficient data flow optimization strategy is adopted to ensure that data can be quickly transmitted between multiple computing units, thereby reducing the delay in the data transmission process. This memory management mechanism effectively alleviates the problem of insufficient memory bandwidth on traditional hardware platforms, enabling the present invention to maintain efficient computing performance in large language model reasoning.

[0041] This embodiment adopts a load-calculate-store periodic process in the calculation process to optimize the calculation throughput of each computing unit. This periodic process reduces the idle time in the calculation process by executing the load, calculate and store operations in parallel. In particular, in each computing unit, the load, calculate and store operations of the computing task can be executed in parallel, ensuring that each computing unit can maximize the use of its computing resources when executing the task. This optimization scheme avoids the situation of idle computing resources in traditional methods, and ensures the maximum utilization of hardware resources by accurately controlling the working state of each computing unit, thereby improving the overall efficiency of the calculation.

[0042] The high-performance large language model (LLM) accelerator based on FPGA of this embodiment can effectively solve the main problems mentioned in the background technology by virtue of its unique hardware architecture design and optimized data flow management solution, especially in terms of computing efficiency, memory bandwidth bottleneck, latency and energy efficiency, and has significant technical effects.

[0043] In the prior art, hardware such as GPU often faces the problem of low computing efficiency due to the large scale of computing tasks, especially when processing large language models, the computing resources are not fully utilized. However, the present invention adopts the FPGA platform and takes advantage of its highly programmable characteristics to parallelize the computing tasks of large language models and assign them to multiple computing units (CUs) for processing. Each computing unit performs efficient calculations through a matrix processing unit (MPE). The parallel work of multiple computing units greatly improves the computing efficiency and avoids the waste of computing resources, thereby significantly improving the overall speed of the reasoning process.

[0044] On traditional hardware platforms, especially GPUs, memory bandwidth usually becomes a bottleneck that limits computing performance when processing large-scale computing tasks. However, the present invention optimizes memory management and the transmission path of data streams, and adopts a solution combining high-bandwidth memory (HBM) and off-chip memory (such as DDR) to effectively improve memory access efficiency. Multiple buffers, DMA data transmission mechanism, and data parallel processing structure in the memory management module effectively solve the bottleneck problem of data transmission during large language model reasoning.

[0045] This embodiment adopts a parallelization strategy of load-compute-store cycles to avoid the problem of waiting for computing tasks in traditional computing methods. By reducing idle time and improving resource utilization of computing units, not only latency is reduced, but also overall energy efficiency is improved. In addition, the customized hardware design of FPGA can flexibly adjust computing resources and memory usage according to the needs of different tasks, thereby achieving more efficient energy management, which is more energy efficient than traditional hardware such as GPU.

[0046] In summary, the FPGA-based large language model accelerator of this embodiment successfully solves the problems of computing efficiency, memory bandwidth bottleneck, latency and energy efficiency mentioned in the background technology by optimizing the parallel computing capability, memory management strategy and data flow path of the computing unit, and provides an efficient and energy-saving solution for large language model reasoning.

[0047] In this embodiment, the matrix processing unit includes multiple configurable DSP cores, which are configured through HLS technology to achieve efficient parallel processing of matrix operations required in large language models.

[0048] In this embodiment, each computing unit in the accelerator communicates with the high-bandwidth memory through an independent AXI interface, and each computing unit can flexibly schedule the memory bandwidth to optimize the utilization of computing resources.

[0049] In this embodiment, the memory management module of the accelerator includes multiple read buffers and write buffers, supports multi-channel parallel access to memory, and implements parallel data transmission through DMA channels to reduce data transmission delay.

[0050] In this embodiment, the accelerator reduces the idle time in the load-compute-store cycle through an optimized data flow structure, ensuring that each computing unit operates efficiently, thereby maximizing the computing throughput.

[0051] In this embodiment, the memory management module of the accelerator dynamically allocates data storage between on-chip memory and off-chip memory according to data flow and computing requirements.

[0052] In this embodiment, the accelerator optimizes the data transmission path to ensure that data is quickly and efficiently transmitted from the host CPU to the FPGA, thereby reducing the overall delay of the system.

[0053] In this embodiment, the accelerator adopts an FPGA platform and is implemented on Xilinx Alveo U280 hardware, optimizing the reasoning process of a large language model through its programmable hardware resources.

[0054] In this embodiment, the accelerator can adjust the configuration of the computing unit according to the characteristics of the large language model to adapt to different types of reasoning tasks and achieve hardware-level acceleration optimization.

[0055] In one embodiment, the rapid development of artificial intelligence has given rise to a variety of network model algorithms. Today, the advancement of artificial intelligence has ushered in an era dominated by large models (Large Language Model, LLM), such as GPT-3.5, GPT-4.0, Llama, Jurassic-1, BERT Turbo, etc. These models can not only understand user input, but also respond quickly to user input, which has revolutionized many fields such as speech, images, and videos. The ability of large models to process and generate data is unlimited, and they play a key role in many high-performance demand scenarios. Tinyllama is a compressed and optimized version of the Llama large model, specifically designed to maintain high accuracy while significantly reducing the model size. When deployed in scenarios such as edge servers, IoT devices, satellite communications, etc., Tinyllama needs to increase computing power to effectively process various data, so as to solve the balance between performance and resource usage, improve performance and reduce energy consumption when deploying AI applications on a large scale.

[0056] The Tinyllama architecture makes specific modifications to the Llama model to reduce its size and improve its running efficiency, making it suitable for environments with limited computing resources. The model architecture retains the basic structure of the Transformer. The LLM predicts the next token based on the contextual information of a given input text sequence, so only the decoder part is retained. The Tinyllama model structure is basically the same as the Transformer Decoder part, mainly consisting of 32 TransformerBlocks, each of which has a structure like Figure 2 shown.

[0057] The input data of LLM is usually one or more paragraphs of natural language text, which can be a simple sentence or a paragraph. The text is represented as a sequence of words or characters to form a token sequence. The token sequence is further serialized into a list or array and indexed through the corpus, mapping each token to a unique integer index for easy calculation within the model. The text information is then converted into a sequence of tokens in digital form, and then the digital tokens are mapped to a real number vector EmbeddingVector through the Embedding layer. Among them, the vector corresponding to each token usually has a fixed dimension d (such as 50, 100, 300, 768, etc.), and each element (real number) in the vector represents a certain attribute or feature of the token in a specific semantic space.

[0058] The Transformer model usually only performs position encoding once after the input sequence passes through the Embedding layer, while the Tinyllama model chooses to perform rotational position encoding (Rotary Positional Embedding, RoPE) on the Query (Q) and Key (K) in each Attention layer. That is, each time the Attention is calculated, the Q and K of the current layer need to be positionally encoded. Then the corresponding rotational position encoding is calculated for each token position; the elements of the Q and K vectors at each token position are rotated in pairs. Specifically, each group of continuous elements of the vector is regarded as a complex number (real and imaginary parts), and then the complex number is rotated according to the rotation angle of the position; the inner product operation is performed on the corresponding elements of each Q vector and all K vectors to obtain the attention score.

[0059] The specific formula of Attention is as follows:

[0060]

[0061] Where Q, K and V represent query, key and value, d kis the dimension of the key. The attention mechanism is the core of the model architecture and helps to handle dependencies regardless of the distance of the input data in the sequence. When given the same set of queries, keys, and values, the model can learn different behaviors based on the same attention mechanism, and then combine the different behaviors as knowledge to capture various ranges of dependencies in the sequence (for example, short-distance dependencies and long-distance dependencies). Therefore, instead of using only a single attention, we can transform queries, keys, and values ​​using h sets (generally h=8) of different linear projections learned independently. Then, these h sets of transformed queries, keys, and values ​​are sent to the attention in parallel. Finally, the outputs of these h attentions are spliced ​​together and transformed by another linear projection that can be learned to produce the final output. This design is called multihead attention, and the formula is as follows:

[0062] MultiHead(Q,K,V)=Concat(head1,...,head h )W O ;

[0063] in

[0064] In this equation, W O , and are parameters learned during training, each corresponding to a different linear transformation of the query, key, and value. These transformations enable the Tinyllama model to operate efficiently.

[0065] After the Attention layer, the FeedForward layer is processed. Tinyllama uses the SwiGLU (SiLU) activation function, and the formula is as follows.

[0066]

[0067] Tinyllama leverages the core technical principles of the Llama model to significantly reduce model size and increase computing requirements while maintaining high accuracy. When deployed in specific scenarios, the Tinyllama model needs to be applied and run at high speed. The use of this patented technology can solve the problem of accelerating the application of large models in high-performance application scenarios and reduce energy consumption when deploying AI applications on a large scale. This not only ensures the wider use of large models, but also breaks through the limitations of current AI hardware technology.

[0068] Although large models are powerful, they also bring significant challenges in terms of data scale and computational requirements. For example, the GPT-4 large model contains hundreds of billions of parameters and requires a lot of memory and computational overhead - more than 1PetaFLOP / s per inference. This has become a major bottleneck for deploying large models in latency-sensitive and resource-constrained environments. In addition, model compression techniques such as sparsification and quantization, although beneficial, often lack support from traditional hardware such as GPUs, especially when dealing with unstructured sparsity. Although sparsity can maintain algorithm accuracy, it cannot be translated into actual performance improvements.

[0069] Given the limitations of current large model hardware, FPGAs are a particularly effective solution. Compared to GPUs, FPGAs have the advantage of flexible model-based hardware customization to better adapt to the unique computing paradigms of large models, such as different sparsity patterns and mixed precision quantization. The reconfigurability of FPGAs allows hardware algorithms to be adjusted to optimize computational throughput and memory utilization, which is critical in large models.

[0070] Deploying large models through hardware accelerators involves several key challenges that affect efficiency and performance. The following are the main challenges encountered in hardware acceleration of large models:

[0071] Computational efficiency and energy consumption: The biggest challenges facing large models in terms of computational efficiency stem from their sheer size and complexity. These models can contain billions to trillions of parameters, and therefore require a lot of computational resources both during training and inference. The density and depth of neural networks in such models means that operations require a large number of matrix multiplications and high data throughput, which puts pressure on traditional hardware systems.

[0072] Memory bandwidth and memory capacity: Large models, especially deep learning models such as neural networks, need to process a large amount of data simultaneously. This places huge demands on the memory bandwidth and memory capacity of the hardware. Accelerators such as GPUs and FPGAs often encounter latency issues caused by transferring large amounts of data between the processor and memory. Overcoming these limitations requires sophisticated memory management techniques to increase throughput and reduce bottlenecks.

[0073] These challenges highlight the complexity of deploying large models in accelerated hardware environments. Addressing these issues often requires multiple optimization methods. FPGAs provide a versatile and energy-efficient platform that can be highly customized to meet the unique computational and memory requirements of modern LLMs. Unlike GPUs, FPGAs can be programmed to support flexible modes and more effectively utilize heterogeneous memory systems, thereby improving computational efficiency and reducing memory bandwidth waste.

[0074] In this embodiment, a high-performance LLM accelerator based on FPGA is designed. This embodiment uses the parallel computing power of FPGA and the general computing power of CPU to provide a complete solution for accelerating large models. Figure 1 The overall structure of the accelerator is given. The accelerator architecture runs on the Xilinx Alevo U280 device, which can effectively accelerate LLM and improve the efficiency of reasoning. The accelerator uses multi-core parallelism to complete LLM reasoning tasks, assigns tasks to different cores through PCIe, and controls data synchronization.

[0075] The architecture of this embodiment is connected by the CPU host and the FPGA device through PCIe. The CPU generates the Tinyllama deployment task and transfers the vector data to the FPGA. According to the tasks in the matrix processing unit (MPE) and the nonlinear operations in the SFU, six cores are divided in the host. These cores run according to the load-compute-store cycle in the hls stream. The FPGA device-side accelerator is based on the Tinyllama design, and its unique architecture provides high parallelism and efficient off-chip memory access. The memory management module in the FPGA accelerator consists of six buffers (including read buffers and write buffers) to support Figure 2 The operation process of Tinyllama algorithm.

[0076] The Tinyllama kernel runs on the host CPU, which sends data to the FPGA HBM via PCIe. Once the data enters the HBM, the MPE can obtain it in parallel through the AXI channels. Since each HBM interface is 256 bits wide, the channel can be divided into multiple channels according to the data bit width. Each channel can be applied to an independent accelerator, called a kernel. The MPE is composed of multiple kernels with multiple communication modes, depending on the high-level description of the functions to be implemented. In order to manage the data exchange between the HBM interface and the parallel kernels, the MPE uses two hardware modules, namely Figure 1 The read and write modules in the MPE are executed in parallel with the kernel and communicate with the data flow model. The read module reads the input data of one element from the HBM memory into the internal buffer, while the write module transfers the corresponding result to the HBM memory through the memory control. To avoid conflicts and reduce the complexity of the AXI interface, we assign these modules to independent HBM channels. The data read and write modules within the MPE split the 256-bit HBM data into 32-bit or 64-bit data for use by the computing module according to the data format.

[0077] Data parallelism can maximize data throughput while minimizing data processing latency and overhead. By improving the way data flows in the system, the overall performance and speed of large model reasoning processes can be improved. A parallel matrix processing unit (MPE) is designed in this accelerator, such as Figure 3 As shown in Figure 1, various stages of data processing can be processed simultaneously, significantly improving throughput. The Direct Memory Access (DMA) channel allows asynchronous and parallel data transfer directly from external memory to the FPGA, bypassing the slower memory bus and CPU. When deploying the Tinyllama model, the hardware platform uses a load-compute-store cycle. In order to minimize the load-compute-store cycle, Figure 3 A parallel matrix processing unit MPE is designed in this paper, which consists of multiple hardware blocks. Each block has multiple DSP48 cores cascaded in a fixed way. The CU is composed of DSP chains, and the MPE is composed of multiple CUs. Each CU has two DSP48 cores. We use configurable cascaded DSPs to support matrix operations and add parallel methods to the DSP chains that need to be accelerated.

[0078] At the same time, applying HLS (High-Level Synthesis) promotes efficient serialization of data, thereby simplifying processing and reducing the need for large memory buffers. This approach is critical for handling fast, sequential access patterns and can reduce the latency associated with random memory access operations.

[0079] The total execution time of Tinyllama can be expressed as T total , the formula is as follows:

[0080] T total =max(T load ,T compute ,T store );

[0081] Where T load 、T compute and T store It represents the time for loading, computing, and storing. The idle time in the FPGA should be minimized, and the time occupied by loading, computing, and storing should be reduced as much as possible.

[0082] T compute The formula is as follows:

[0083]

[0084] Among them, α represents the number of channels for data transmission, β represents the size ratio of intermediate data, and D in Indicates the input data size, #PE lRepresents the number of concurrent processing elements (PEs). Ⅱ represents the startup interval, which is the number of cycles required for the FPGA to process an input of the size of datawidth, indicating the computational intensity of the kernel.

[0085] In T compute In the equation, HLS increases the parallelism of the processing unit and reduces the startup interval. Therefore, we first reduce the number of load-compute-store cycles. Combine the six cores in parallel. In the original method, each core is a loop, and all loops are performed sequentially. Therefore, for Tinyllama, the load-compute-store loop will run hundreds of times. We use HLS to optimize the entire process in the loop, and then calculate Tinyllama in this process. Finally, store them in the host memory.

[0086] In order to optimize the memory, this embodiment implements a hybrid memory method that can fully utilize the advantages of on-chip and off-chip memory, such as Figure 4 , which can maximize the access speed of critical operations while maintaining sufficient memory capacity.

[0087] The HBM memory consists of many independently accessible channels, supports concurrent access between HBMs, and has a wide interface in each channel, allowing multiple data to be read / written from the off-chip memory per clock cycle. In order to avoid switch congestion, it is important to divide the data into different HBM memory areas so that each CU uses as few channels as possible and shares these channels with as few other CUs as possible. Therefore, the present invention proposes an HBM data allocation structure that allocates memory space in advance according to the core size, such as Figure 4 as shown in .

[0088] The U280 card uses the XCU280 FPGA, which contains three Super Logic Regions (SLRs). The SLR is a physical part of the FPGA with a specific number of resources, as shown in Table 1. SLR0 integrates an HBM controller that interacts with the HBM2 subsystem through 32 pseudo channels (PCs), each of which can directly access 256MB of storage space (8GB in total). Each 256-bit PC runs at 450MHz with a maximum bandwidth of 14.4GB / s. Therefore, the entire system can achieve a theoretical bandwidth of 460.8GB / s. SLR0 is also connected to the host through 16 channels of the PCIe interface. SLR0 and SLR1 are each connected to 16GB of DDR4. Finally, each region has up to 8MB of programmable logic RAM (PLRAM) for fast access to small data sets.

[0089] The target system of Alveo U280 consists of multiple computing units (CUs). Each CU is a user-defined hardware module that can be connected to any PC through an independent AXI interface, while the built-in HBM controller can access all physical channels. The CU can be described in C++ and synthesized with HLS, or it can be specified directly in RTL. Multiple CUs allow parallel execution but must be connected to different HBM channels. The system configuration file describes the connection between CU ports and HBM channels. The required logic is automatically generated during system synthesis.

[0090] Table 1. FPGA platform hardware parameters

[0091] FPGA Hardware Platform Xilinx Alevo U280 (16nm) frequency 225MHz Computational Unit 9024 DSPs Storage Space 8 & 32GB bandwidth 460&38GB / s

[0092] Throughput: This embodiment compares the throughput of the unoptimized, non-parallelized, non-memory optimized and the method of the present invention. Compared with before optimization, the throughput of the present invention on U280 is improved by 1.6 times, as shown in Table 1. The bandwidth utilization of this accelerator on FPGA is better than that of the unoptimized accelerator. This is because the accelerator can customize the design of hardware units to fully utilize the parallelism and storage optimization of LLM. In contrast, the unoptimized accelerator adopts an iterative design, which can iteratively reuse various operators on hardware units, thereby improving resource utilization.

[0093] Latency: Latency is measured as a time function in the host program. The present invention calculates latency from kernel startup, including FPGA communication time. The latency measurement results of the present accelerator are better than those of the unoptimized accelerator, achieving up to 4.8 times speedup in latency. This is mainly due to the high-performance core based on Tinyllama and the efficient combination of multiple cores in the present accelerator.

[0094] Energy efficiency: This embodiment further evaluates energy efficiency. The power consumption of the U280 device was measured using the Xilinx Board Utility tool xbutil, and the average of five independent measurements was taken, and the final power consumption was determined to be 38.17W. The non-parallel and non-memory optimized ones were 36.24W and 28.39W, respectively. Compared with the non-memory optimized accelerator, the method of the present invention achieved 1.01 times the energy efficiency, which is mainly due to the reduction of redundant off-chip memory communications through the llama model. Compared with the non-parallel accelerator, this method achieved 1.08 times the energy efficiency. Parallel processing can save a lot of unnecessary time. Compared with non-optimization, the energy efficiency of the present invention is 1.18 times higher than that of the non-optimized accelerator.

[0095] Table 2. Comparison of acceleration effects

[0096]

[0097] In one embodiment, a method for accelerating large language model reasoning based on FPGA is provided, wherein the method comprises the following steps:

[0098] Generate large language model inference tasks in the host CPU and transfer the data to the FPGA through the PCIe interface;

[0099] Using multiple computing units in the FPGA to perform inference tasks in parallel, each computing unit includes a matrix processing unit;

[0100] Use high-bandwidth memory to provide data to the computing unit and read and write data efficiently through parallel memory access;

[0101] In the load-compute-store cycle, the computing throughput of each computing unit is optimized and the idle time is reduced.

[0102] This embodiment proposes an FPGA acceleration hardware structure based on the Tinyllama large model, which solves the key problem of low efficiency in traditional FPGA neural network implementation; proposes a data parallel structure based on load-compute-store iterations, which can minimize iterations and time-consuming cycles, and improve throughput and reduce execution time by ensuring that computing units continuously obtain data and avoiding idle time. A memory allocation structure is proposed to achieve the circular use of memory based on Tinyllama, in which each computing unit will be reused after data processing is completed without waiting for all processing to end. This circular reuse is managed through efficient memory scheduling, which improves the efficiency of data input to the processor.

[0103] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A high-performance large language model accelerator based on FPGA, characterized in that: The accelerator comprises: A plurality of computing units, wherein the computing units are connected to a host CPU via a PCIe bus and are connected to a high bandwidth memory; Each of the computing units includes at least one matrix processing unit and performs parallel computing according to a periodic process of loading-computing-storing; The accelerator optimizes the data loading, computing and storage processes through parallel computing among multiple computing units and their memory scheduling mechanism, wherein the multiple computing units work in coordination to ensure the efficient operation of each computing unit when loading data, executing computing tasks and storing results; the memory scheduling mechanism is responsible for dynamically allocating memory resources, optimizing data storage locations and ensuring optimal data access for each computing unit under memory bandwidth constraints.

2. The high-performance large language model accelerator according to claim 1, characterized in that: The matrix processing unit includes multiple configurable DSP cores, which are configured through HLS technology to achieve efficient parallel processing of matrix operations required in large language models.

3. The high-performance large language model accelerator according to claim 1, characterized in that: Each computing unit in the accelerator communicates with the high-bandwidth memory through an independent AXI interface, and each computing unit can flexibly schedule the memory bandwidth to optimize the utilization of computing resources.

4. The high-performance large language model accelerator according to claim 1, characterized in that: The memory management module of the accelerator includes multiple read buffers and write buffers, supports multi-channel parallel access to memory, and realizes parallel transmission of data through DMA channels to reduce data transmission delay.

5. The high-performance large language model accelerator according to claim 1, characterized in that: The accelerator reduces the idle time in the load-compute-store cycle through an optimized data flow structure, ensuring that each computing unit operates efficiently, thereby maximizing computing throughput.

6. The high-performance large language model accelerator according to claim 1, characterized in that: The memory management module of the accelerator dynamically allocates data storage between on-chip memory and off-chip memory according to data flow and computing requirements.

7. The high performance large language model accelerator according to claim 1, characterized in that: The accelerator optimizes the data transmission path to ensure that data is quickly and efficiently transmitted from the host CPU to the FPGA, thereby reducing the overall system latency.

8. The high-performance large language model accelerator according to claim 1, characterized in that: The accelerator adopts an FPGA platform and is implemented on Xilinx Alveo U280 hardware, optimizing the reasoning process of large language models through its programmable hardware resources.

9. The high-performance large language model accelerator according to claim 1, characterized in that: The accelerator can adjust the configuration of the computing unit according to the characteristics of the large language model to adapt to different types of reasoning tasks and achieve hardware-level acceleration optimization.

10. A method for accelerating large language model reasoning based on FPGA, characterized in that: The method comprises the following steps: Generate large language model inference tasks in the host CPU and transfer the data to the FPGA through the PCIe interface; Using multiple computing units in the FPGA to perform inference tasks in parallel, each computing unit includes a matrix processing unit; Use high-bandwidth memory to provide data to the computing unit and read and write data efficiently through parallel memory access; In the load-compute-store cycle, the computing throughput of each computing unit is optimized and the idle time is reduced.

Citation Information

Cited By

  • Large model calculation acceleration chip architecture

    CN120525012A

  • Large model computing acceleration chip

    CN120525012B

  • Heterogeneous system and method for efficiently realizing attention mechanism of large language model and application

    CN121436044A