Compilation optimization device and method supporting Page Attention and medium

By optimizing the Page Attention computation process with a multi-core AI accelerator, we solved the problems of low KV Cache management efficiency and complex multi-batch dynamic scheduling on the ASIC platform, achieved more efficient computing and resource utilization, and improved the inference performance of large language models.

CN120704681APending Publication Date: 2025-09-26BEIJING YIXIN YIYU MICROELECTRONICS TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510661576.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies for Page Attention calculations on large language models suffer from insufficient hardware adaptability, computational efficiency bottlenecks, and poor dynamic sequence support. In particular, they fail to fully utilize the advantages of parallel computing and high bandwidth on ASIC architectures, and fail to deeply optimize graphics memory management and computing logic.

Method used

It uses a multi-core AI accelerator, including tensor components, vector components, DMA components, and IMM components. Through the KV Cache management module, DMA control module, and multi-batch inference scheduling module, it implements standardized storage, dynamic adjustment, data handling optimization, and dynamic allocation of computing resources for KV Cache blocks. It combines double buffering technology and sparse matrix calculation to optimize the Page Attention calculation process.

Benefits of technology

It improves bandwidth utilization, reduces computing latency, enhances dynamic adaptability, maximizes hardware resource utilization, reduces redundant computing, and improves inference efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704681A_ABST
    Figure CN120704681A_ABST
Patent Text Reader

Abstract

The invention discloses a compiling optimization device and method supporting Page Attention and a medium, and belongs to the field of compiling optimization devices.The device comprises a multi-core AI accelerator, and each calculation core comprises a tensor component, a vector component, a DMA component and an IMM component; the compiling optimization device is optimized through the following modules: (a) a KV Cache management module, which is used for storing page table items corresponding to different KV Cache blocks according to a batch sequence and generating standardized page table items; (b) a DMA control module which supports a linear reading and remapping mode, carries discrete KV Cache blocks from an external memory to an on-chip weight buffer area in a segmented manner according to an index table, and hides data loading delay by adopting a double-buffer technology; and (c) a multi-batch reasoning scheduling module which dynamically allocates computing resources according to the data scale in a code stage. According to the invention, through deep cooperation of a hardware architecture and functional logic, the problems of low KV Cache management efficiency, complex multi-batch dynamic scheduling and insufficient computing resource utilization rate on an ASIC platform are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of compilation optimization devices, and in particular to a compilation optimization device, method and medium supporting Page Attention. Background Art

[0002] As large language models (such as GPT and LLaMA) grow in size, they require maintaining a massive KVCache (key-value cache) during inference, leading to a surge in video memory usage, memory fragmentation, and decreased batch processing efficiency. Existing technologies, such as the VLLM framework, alleviate the problem of video memory waste on GPU platforms by managing KV Cache through paging (similar to the operating system's virtual memory mechanism). However, the following drawbacks remain:

[0003] (1) Insufficient hardware adaptation: Existing solutions are mainly optimized for GPUs and lack adaptation to the characteristics of ASIC architecture, and cannot fully utilize its parallel computing and high bandwidth advantages;

[0004] (2) Computational efficiency bottleneck: Only graphics memory management is optimized, without further optimizing the computational logic of Page Attention itself (such as multi-batch scheduling and data movement delay);

[0005] (3) Poor support for dynamic sequences: When the sequence lengths of different batches change dynamically, existing methods are unable to flexibly schedule hardware resources, resulting in a waste of computing resources.

[0006] Therefore, those skilled in the art provide a compilation optimization device, method, and medium supporting Page Attention to solve the problems raised in the above background technology. Summary of the Invention

[0007] In view of the defects in the prior art, the present invention provides a compilation optimization device supporting Page Attention, comprising:

[0008] Multi-core AI accelerator, each computing core includes a tensor component, a vector component, a DMA component, and an IMM component;

[0009] The compilation optimization device realizes optimization through the following modules:

[0010] (a) A KV Cache management module, which stores page table entries corresponding to different KV Cache blocks in batch order and generates standardized page table entries. The content of each page table entry is the offset of the corresponding KV Cache block address relative to the base address, which can be dynamically adjusted.

[0011] (b) DMA control module, which supports linear read and remapping modes, and moves discrete KV cache blocks from external memory to the on-chip weight buffer in segments according to the index table, and uses double buffering technology to hide data loading delays;

[0012] (c) Multi-batch inference scheduling module dynamically allocates computing resources based on data size during the decoding phase:

[0013] When the data size is smaller than the threshold (small-scale data), multiple batches share the weight buffer and are collaboratively calculated by multiple computing cores, and masks are generated by the IMM component to shield redundant data;

[0014] When the data size is greater than or equal to the threshold (large-scale data), different batches are assigned to independent computing cores, and multiple stages of attention calculation are interleaved in a pipeline manner.

[0015] As a further solution of the present invention: in the KV Cache management module, the method for generating a page table entry includes:

[0016] The offsets of the KV Cache block addresses of different batches relative to the base address are divided by the block size to obtain standardized page table entries, and the page table entries are stored continuously in batch order.

[0017] As a further solution of the present invention: the DMA control module further includes:

[0018] Weight prefetch logic, which prefetches the next batch of KV Cache blocks to the backup buffer while the tensor component is processing the current weight data;

[0019] The dynamic scheduling unit adjusts the data transfer order according to the task load to avoid DDR bandwidth resource conflicts.

[0020] As a further solution of the present invention: in the multi-batch inference scheduling module, the mask generation logic is specifically as follows:

[0021] The IMM component generates a mask vector for each computing kernel, which only embeds 0 in the valid interval of the current batch and sets the rest to negative infinity. After the Softmax calculation, the redundant data is reset to zero, adapting to the sparse matrix to accelerate the calculation.

[0022] As a further solution of the present invention: the compilation optimization device performs the following steps in the decoding stage:

[0023] (1) Allocate query matrices of multiple batches to different computing cores;

[0024] (2) Load multiple batches of KV Cache blocks into the weight buffer in parallel through the DMA component;

[0025] (3) The tensor component performs matrix multiplication of query and key, superimposes the causal mask and calculates the softmax;

[0026] (4) The vector component completes the weighted summation with the Value matrix and outputs the Attention result.

[0027] As a further solution of the present invention: in step (3), the generation of the causal mask and the Softmax calculation are completed by the vector component, and the zero element calculation is skipped through the sparse matrix optimization logic.

[0028] The present application also discloses a compilation optimization method supporting Page Attention, which is applied to a compilation optimization device supporting Page Attention, comprising the following steps:

[0029] S1. Design the continuous layout format of the KVCache and generate a standardized page table based on the number of PEs and bandwidth of the target hardware.

[0030] S2. During tensor computation, the DMA component dynamically loads KVCache blocks into the double-buffered weight area, enabling parallelization of data movement and computation.

[0031] S3. In the decoding phase, select multi-batch collaboration or independent core allocation strategies based on the data scale, and combine mask generation with sparse computing optimization to improve inference efficiency.

[0032] As a further solution of the present invention: in step S2, the double buffer weight area is divided into a ping area and a pong area, and the DMA component preloads the next batch of KV Cache blocks into the pong area when the tensor component reads the data in the ping area.

[0033] As a further solution of the present invention: in step S3, the multi-batch collaboration strategy specifically includes:

[0034] The KV Cache blocks of multiple batches are spliced ​​into the same weight buffer. Multiple computing cores share data and perform parallel calculations. The mask vector only retains the valid data of the current batch.

[0035] The present application also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a compilation optimization method supporting Page Attention.

[0036] The beneficial effects of the present invention are embodied in:

[0037] 1. Improved bandwidth utilization: By continuously storing page table entries and coordinating multi-batch reads, bandwidth waste caused by dynamic address access is reduced.

[0038] 2. Reduced computing latency: Double buffering technology and DMA prefetch logic hide data loading latency and improve PE computing efficiency;

[0039] 3. Enhanced dynamic adaptability: Mask generation and sparse computation optimization support dynamic changes in batch sequence lengths, reducing redundant computations.

[0040] 4. Efficient utilization of hardware resources: The pipelined multi-core scheduling strategy maximizes hardware parallelism and avoids idle resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly describes the drawings required for the specific embodiments or the description of the prior art. Similar elements or parts are generally identified by similar reference numerals throughout the drawings. Elements or parts in the drawings are not necessarily drawn to scale.

[0042] Figure 1 The hardware structure block diagram of a compilation optimization device supporting Page Attention;

[0043] Figure 2 This paper presents the intent of a KV Cache page in a compilation optimization device that supports Page Attention.

[0044] Figure 3 A computational graph of tensor components in a compilation and optimization device supporting Page Attention;

[0045] Figure 4 A multi-batch KV cache splicing diagram in a compilation optimization device supporting Page Attention;

[0046] Figure 5 Schematic diagram of the execution process of multiple batches on different cores in a compilation optimization device that supports Page Attention. DETAILED DESCRIPTION

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0048] As mentioned in the background technology of this application, research has found that the existing Page Attention technology alleviates the problem of video memory waste on GPU platforms (such as the VLLM framework) by managing KV Cache by paging (similar to the virtual memory mechanism of the operating system), but it still has the following defects: (1) Insufficient hardware adaptation; (2) Computational efficiency bottleneck; (3) Poor dynamic sequence support.

[0049] To address the above-mentioned shortcomings, this application discloses a compilation optimization device, method, and medium that supports Page Attention. Through deep collaboration between hardware architecture and functional logic, it solves the problems of low KV Cache management efficiency, complex multi-batch dynamic scheduling, and insufficient computing resource utilization on ASIC platforms.

[0050] The following will describe in detail how the solution of this application solves the above technical problems with reference to the accompanying drawings.

[0051] See also Figure 1 In an embodiment of the present invention, a compilation optimization device supporting Page Attention includes: a multi-core AI accelerator, each computing core includes a tensor component, a vector component, a DMA component, and an IMM component; wherein the tensor component implements tensor calculations of operators such as MatMul, the vector component is used for vector calculations of operators such as Mul, Add, and Softmax, the DMA component implements data movement from DDR to on-chip space, and the IMM component is used to generate a mask matrix. During calculation, the input data is first moved from the DDR to the L2 cache, and the tensor component and the vector component load the required local data blocks from the L2 cache and store them in the L1 cache to accelerate the calculation. The L1 cache provides operands for PE, supporting high-speed calculations.

[0052] The compilation optimization device achieves optimization through the following modules: (a) KV Cache management module, which is used to store the page table entries corresponding to different KV Cache blocks in batch order and generate standardized page table entries. The content of the page table entry is the offset of the corresponding KV Cache block address relative to the base address, which can be dynamically adjusted; (b) DMA control module, which supports linear read and remapping mode, and moves discrete KV Cache blocks from external memory to the on-chip weight buffer in segments according to the index table, and uses double buffering technology to hide data loading delay; (c) Multi-batch inference scheduling module, which dynamically allocates computing resources according to data size in the decoding stage: when the data size is less than the threshold (small-scale data), multiple batches share the weight buffer and are collaboratively calculated by multiple computing cores, and masks are generated by the IMM component to shield redundant data; when the data size is greater than or equal to the threshold (large-scale data), different batches are assigned to independent computing cores, and multiple stages of attention calculation are interleaved in a pipeline manner. The three modules (KV Cache management module, DMA control module, and multi-batch inference scheduling module) are not independent hardware components, but functional logic built based on the hardware components of the multi-core AI accelerator. Their specific performance is as follows:

[0053] KV Cache management module: This module is implemented by combining the address remapping capability of the DMA component and the calculation logic of the tensor component with the continuous storage design of page table entries.

[0054] DMA control module: relies on the double buffering mechanism, prefetch logic and dynamic scheduling capabilities of the DMA component to optimize data loading efficiency;

[0055] Multi-batch inference scheduling module: This module leverages the parallelism of multiple computing cores (such as collaborative computing of multiple tensor components) and combines it with the mask generation function of the IMM component to achieve dynamic resource allocation.

[0056] In this embodiment, in the KV Cache management module, the method for generating page table entries includes: dividing the offset of the KVCache block address of different batches relative to the base address by the block size to obtain a standardized page table entry, and continuously storing the page table entries in batch order. First, the present invention designs an optimized KVCache data arrangement scheme based on the number of PEs, the number of computing cores and the bandwidth of this application. Secondly, as Figure 2As shown, the Key Cache blocks of different batches have different addresses. For example, the address of the first Key Cache block of batch 0 is 0x2000, and the address of the first Key Cache block of batch 1 is 0x1000. The offset of the address of each Key Cache block relative to the zero address is divided by the size of the Key Cache block to obtain a page table entry. For example, the address of the first Key Cache block of batch 0 is 0x2000, and the block size is 0x1000, so its corresponding page table entry is 2. In order to improve the continuity of reading when processing multiple batches of data, the present invention stores the page table entries corresponding to different KV Cache blocks in the order of batches.

[0057] In this embodiment, the DMA control module further includes: weight prefetch logic, which prefetches the next batch of KV Cache blocks to the backup buffer when the tensor component processes the current weight data; a dynamic scheduling unit, which adjusts the data transfer order according to the task load to avoid DDR bandwidth resource conflicts. In order to speed up the reading speed of weight data and avoid it becoming a bottleneck for tensor calculation, this application designs a dedicated weight buffer. Figure 3 As shown in the figure, each time a tensor calculation is performed, the tensor component sends a request to the DMA component through the table lookup logic, automatically loading multiple non-contiguous KV Cache blocks from the DDR into the weight buffer. This process is completed by reading the KV Cache page table. During the calculation process, the tensor component has a weight data prefetch function: when processing the current weight data, the prefetch logic will pre-load the next batch of weights into the buffer, effectively hiding the data loading delay. Figure 3In [1], the weight buffer is divided into two areas: ping and pong. While the tensor component reads KVCache data from the ping area and feeds it into the PE array for computation, the DMA component can pre-load some KV Cache into the pong area. Furthermore, this design dynamically adjusts the timing and order of weight migration based on the actual task execution, enabling optimal scheduling of DDR bandwidth, avoiding resource conflicts, and improving overall efficiency. During Attention calculations, the Key and Value matrices are typically of the shape (batch, head_num, seq_len, head_dim), which serve as the weight matrices for Matmul calculations. However, due to the characteristics of Matmul and Attention calculations, the Key and Value matrices often need to be transposed to ensure correct results. The DMA component of this application supports two methods for loading weight data: linear read and remapping. The read process is split into two sides: the source side and the destination side. The source side is responsible for reading weight data from the DDR, while the destination side writes the data to the weight buffer. Both sides support linear read and remapping. Linear read maximizes bandwidth utilization, while remapping allows for transposition of weight data, providing greater flexibility. By combining these two methods, we can optimize bandwidth utilization and improve computing efficiency to meet different computing needs.

[0058] In this embodiment, in the multi-batch inference scheduling module, the mask generation logic is specifically as follows: the IMM component generates a mask vector for each computing core, only setting 0 in the current batch valid interval and setting the rest to negative infinity. After the Softmax calculation, the redundant data is reset to zero, and the sparse matrix is ​​adapted to accelerate the calculation. This application uses the DMA component to read the KCache blocks of multiple batches from the DDR and store them in the on-chip weight cache area. Figure 4 As shown, when the amount of data is small, the tensor component can be read by Figure 2 The page table in the Query Cache is used to load 4 batches of K Cache blocks at one time. In the decoding stage, the length of the Query matrix is ​​always 1, so the present invention will distribute the Query matrix of 4 batches to 4 computing cores, each core shares the weight data and performs tensor calculations. Since multiple cores share the weight area, there is some redundant calculation in each core. However, in large language models, the product of the Query and Key matrices needs to be added with a causal mask and calculated by Softmax, in which the negative infinite part will eventually return to zero. Therefore, the present invention uses the IMM component on the target device to generate the following for each core: Figure 4Specifically, for a certain kernel, its mask vector is 0 only within the valid range of the current batch, and is negative infinity for the rest of the time. For example, the weight buffer is arranged in the order of batches 0, 1, 2, and 3, and the L2 cache of kernel 0 stores the query matrix data for batches 0, 1, 2, and 3. When calculating batch 0, kernel 0 is only valid for the front-end region. Therefore, kernel 0 is only valid for the front-end region, and the rest of the time is masked to negative infinity. Although this method increases the computational load of each kernel, considering that the target device has high computing power and the decoding stage is limited by bandwidth bottlenecks, this additional computation has little impact on overall performance when the data volume is small. Conversely, the dynamic nature of multi-batch processing will incur additional hardware overhead. On the one hand, the KV Cache page table entry addresses of different batches are dynamic; on the other hand, the sequence lengths of different batches vary, and the called kernels also vary, resulting in increased hardware scheduling complexity. After adopting the method of the present invention, the performance of the target device in the decoding stage will be significantly optimized. By adopting this approach, the Softmax output on each core is a sparse matrix. The target device has special acceleration optimizations for sparse matrix multiplication, thus avoiding invalid calculations on zero elements. The target tensor component has a good acceleration effect on sparse matrices.

[0059] In this embodiment, the compilation optimization device performs the following steps in the decode stage: (1) assigning multiple batches of Query matrices to different computing cores; (2) loading multiple batches of KV Cache blocks to the weight buffer in parallel through the DMA component; (3) the tensor component performs matrix multiplication of Query and Key, superimposes the causal mask and calculates Softmax; (4) the vector component completes the weighted summation with the Value matrix and outputs the Attention result. When processing large-scale data, this solution divides the Attention calculation into multiple stages to improve computational efficiency. Figure 5As shown, these stages include reading the Query matrix, reading the Key matrix, tensor calculation of Query and Key, generation of causal mask, Softmax calculation, reading the Value matrix, tensor calculation of Value, and writing back the final result. Since DMA can access the L2 cache of different cores at the same time, in the decode stage, the length of the Query matrix is ​​always 1, so the 4 batches of Queries can be directly assigned to the 4 cores for reading. Each batch is assigned to an independent core, and different stages of the inference process are executed in an interleaved manner. For example, when batch0 uses the vector component for Softmax calculation on core0, batch1 uses the tensor component for the first round of tensor calculation on core1, and batch2 uses the DMA component to read the Key Cache on core2. This pipelined task division can make full use of hardware resources and further improve the inference efficiency of the device.

[0060] In this embodiment, in step (3), the generation of the causal mask and the Softmax calculation are completed by the vector component, and the zero element calculation is skipped through the sparse matrix optimization logic.

[0061] This application also discloses a compilation optimization method supporting Page Attention, which is applied to a compilation optimization device supporting Page Attention. The method comprises the following steps: S1. Designing a continuous layout format for the KVCache based on the number of PEs and bandwidth of the target hardware and generating a standardized page table; S2. Dynamically loading KVCache blocks into the double-buffered weight area via the DMA component during tensor computation to parallelize data movement and computation; S3. During the decoding phase, selecting a multi-batch collaborative or independent core allocation strategy based on the data size, combined with mask generation and sparse computation optimization, improves inference efficiency. Existing solutions primarily focus on efficient management of KV Cache memory on GPUs. While optimizing the kernel for KV Cache reads, they do not specifically design efficient computational solutions for page attention. This solution optimizes multi-batch page attention based on the common use of multi-batch computation in the decoding phase. In this solution, the multi-batch KV Cache page tables are stored contiguously, enabling more efficient KV Cache reads from the target device and reducing performance degradation caused by dynamic shape changes. At the same time, the target device has the ability to automatically move discrete weight data to on-chip space. This process is performed in parallel with tensor calculations, thereby improving overall computing efficiency. In addition, during the process of moving weight data, the hardware can also convert it according to the arrangement of KV Cache data, further enhancing the flexibility of tensor calculations. Finally, this solution not only uses multi-core parallel computing to improve bandwidth utilization and computing efficiency when the data volume is small, but also splits the Attention calculation into multiple stages when the data volume is large, and uses a pipelined approach to execute different calculation stages of different batches to minimize hardware resource waste.

[0062] In this embodiment, in step S2, the double buffer weight area is divided into a ping area and a pong area. When the tensor component reads data in the ping area, the DMA component preloads the next batch of KV Cache blocks into the pong area.

[0063] In this embodiment, in step S3, the multi-batch collaboration strategy specifically includes: splicing the KV Cache blocks of multiple batches into the same weight buffer, having multiple computing cores share data and perform parallel computing, and the mask vector only retains the valid data of the current batch.

[0064] The present application also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a compilation optimization method supporting Page Attention.

[0065] To further illustrate the present invention, a compilation optimization device, method, and medium supporting PageAttention provided by the present invention are described in detail below with reference to embodiments.

[0066] Example 1: KV Cache Continuous Arrangement and Page Table Generation

[0067] 1.1 Hardware Configuration: The target device is a 4-core AI accelerator, each core of which contains a tensor component (supporting sparse matrix acceleration) and a DMA component (supporting linear / remapping modes);

[0068] 1.2 Page table generation:

[0069] Define the base address as 0x1000 and the KV Cache block size as 0x1000;

[0070] The key cache block address of batch 0 is 0x2000, and the page table entry is (0x2000-0x1000) / 0x1000=1;

[0071] Each batch page table entry is stored sequentially and continuously to ensure that DMA can read in batches;

[0072] 1.3 Effect: When accessing a KV Cache of 4 batches, the hardware overhead is reduced by 30%.

[0073] Example 2: Multi-batch collaborative computing optimization

[0074] 2.1 Small-scale data scenario:

[0075] Concatenate the KV Cache blocks of 4 batches into the same weight buffer;

[0076] The IMM component generates a mask vector for each core (the valid interval is set to 0 and the rest is set to negative infinity);

[0077] After the vector component performs Softmax, the redundant data is reset to zero, and the tensor component skips the calculation of zero elements;

[0078] 2.2 Effect: Bandwidth utilization increased by 40% and inference latency reduced by 25%.

[0079] Example 3: Large-scale data pipeline scheduling

[0080] 3.1 Hardware allocation: Allocate 4 batches to 4 independent cores;

[0081] 3.2 Stage Division:

[0082] Core 0: Reads the Query matrix and executes MatMul;

[0083] Core 1: Load the Key Cache and generate the causal mask;

[0084] Core 2: Calculate the weighted sum of Softmax and Value;

[0085] Core 3: writes back the final result;

[0086] 3.3 Effect: Hardware resource utilization reaches 90%, and throughput increases by 3 times.

[0087] It should be noted that the following explanations are made for the terms in the above content:

[0088] MatMul: Matrix Multiplication;

[0089] Input sequence: an input matrix representing the word embedding or feature vector of the sequence, which is the basis for self-attention calculation on the input;

[0090] Query matrix: The query matrix is ​​obtained by linear transformation of the input sequence and is used to calculate the correlation between the current element and other elements;

[0091] Key: The key is obtained by linear transformation of the input sequence and is used to determine the characteristics and positions of other elements;

[0092] Value: The value is obtained by linear transformation of the input sequence, which represents the specific content or information of each element;

[0093] Token: In natural language processing, a token is the smallest unit of input data, usually the basic unit formed after text is segmented. The specific form depends on the task requirements. A token can be a word, a character, or even a symbol or special mark;

[0094] KV Cache: In autoregressive tasks, each time a new token is generated, the keys and values ​​of the entire context must be accessed. To avoid repeated computations, the keys and values ​​of previous time steps are cached, improving computational efficiency.

[0095] Page Attention: This attention algorithm is inspired by the classic concepts of virtual memory and paging in operating systems. It divides the KV cache of each sequence into blocks, each containing a fixed number of token keys and values. This effectively solves the memory bottleneck problem encountered by large language models when generating text.

[0096] VLLM: An inference and serving engine optimized for high throughput and memory efficiency for large language models.

[0097] GPU: Graphics Processing Unit, a processor dedicated to processing graphics and image calculations;

[0098] ASIC: Application-Specific Integrated Circuit, is an integrated circuit customized for a specific application or task;

[0099] Kernel: A function that implements some core computing logic and can be executed on various accelerators.

[0100] Batch: refers to a group of text samples processed at one time during model training or inference;

[0101] Prefill: refers to the process in which the model calculates hidden states, attention weights, and initial probability distributions based on the input sequence (such as text or other context) in the generation task, providing a basis for the subsequent decoding stage;

[0102] Decode: This refers to the process by which the generative model gradually generates the target sequence, triggered by the prefill phase. The decoding phase is usually performed step by step, with each step inferring the next token based on the currently generated partial sequence until the termination condition is met.

[0103] PE: Processing Element, the basic unit of the computing array, used to perform accelerated computing;

[0104] Online data: refers to data generated in real time during model operation, which is used for real-time training or updating of models;

[0105] Softmax: A commonly used activation function, mainly used in the output layer of multi-class classification problems. It converts a set of arbitrary real numbers into numerical values ​​representing a probability distribution, so that each number is between 0 and 1 and the sum of all numbers is 1;

[0106] DMA component: Direct Memory Access, direct memory access component;

[0107] IMM component: Immediate Operand Unit, immediate operation component;

[0108] Mul: multiplication operator; Add: addition operator; Softmax: normalized exponential function;

[0109] DDR: Double Data Rate Synchronous Dynamic Random Access Memory, double data rate synchronous dynamic random access memory, is a widely used computer memory technology;

[0110] Double Buffer: This is an optimization technique widely used in computer system and hardware design. By setting up two buffers, data generation and consumption can be carried out in parallel, thereby reducing waiting time and improving system efficiency.

[0111] This invention reduces bandwidth waste caused by dynamic address access by continuously storing page table entries and coordinating multi-batch reads. Double buffering technology and DMA prefetching logic conceal data loading delays, improving PE computing efficiency. Mask generation and sparse computation optimization support dynamic changes in batch sequence lengths, reducing redundant computations. A pipelined multi-core scheduling strategy maximizes hardware parallelism and avoids idle resources.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and description of the present invention.

Claims

1. A compilation optimization device supporting Page Attention, characterized in that: include: Multi-core AI accelerator, each computing core includes a tensor component, a vector component, a DMA component, and an IMM component; The compilation optimization device realizes optimization through the following modules: (a) A KV Cache management module, which stores page table entries corresponding to different KV Cache blocks in batch order and generates standardized page table entries. The content of each page table entry is the offset of the corresponding KV Cache block address relative to the base address, which can be dynamically adjusted. (b) DMA control module, which supports linear read and remapping modes, and moves discrete KV cache blocks from external memory to the on-chip weight buffer in segments according to the index table, and uses double buffering technology to hide data loading delays; (c) Multi-batch inference scheduling module dynamically allocates computing resources based on data size during the decoding phase: When the data size is smaller than the threshold, multiple batches share the weight buffer and are collaboratively calculated by multiple computing cores, and masks are generated by the IMM component to shield redundant data; When the data size is greater than or equal to the threshold, different batches are assigned to independent computing cores, and multiple stages of attention calculation are interleaved in a pipeline manner.

2. The compilation optimization device supporting Page Attention according to claim 1, characterized in that: In the KV Cache management module, the method for generating a page table entry includes: The offsets of the KV Cache block addresses of different batches relative to the base address are divided by the block size to obtain standardized page table entries, and the page table entries are stored continuously in batch order.

3. The compilation optimization device supporting Page Attention according to claim 2, characterized in that: The DMA control module further includes: Weight prefetch logic, which prefetches the next batch of KV Cache blocks to the backup buffer while the tensor component is processing the current weight data; The dynamic scheduling unit adjusts the data transfer order according to the task load to avoid DDR bandwidth resource conflicts.

4. The compilation optimization device supporting Page Attention according to claim 3, characterized in that: In the multi-batch inference scheduling module, the mask generation logic is specifically as follows: The IMM component generates a mask vector for each computing kernel, which only embeds 0 in the valid interval of the current batch and sets the rest to negative infinity. After the Softmax calculation, the redundant data is reset to zero, adapting to the sparse matrix to accelerate the calculation.

5. The compilation optimization device supporting Page Attention according to claim 4, characterized in that: The compilation optimization device performs the following steps in the decode stage: (1) Allocate query matrices of multiple batches to different computing cores; (2) Load multiple batches of KV Cache blocks into the weight buffer in parallel through the DMA component; (3) The tensor component performs matrix multiplication of query and key, superimposes the causal mask and calculates the softmax; (4) The vector component completes the weighted summation with the Value matrix and outputs the Attention result.

6. The compilation optimization device supporting Page Attention according to claim 5, characterized in that: In step (3), the generation of the causal mask and the Softmax calculation are completed by the vector component, and the zero element calculation is skipped through the sparse matrix optimization logic.

7. A compilation optimization method supporting Page Attention, characterized in that: The compilation optimization device supporting Page Attention as claimed in any one of claims 1 to 6 comprises the following steps: S1. Design the continuous layout format of the KVCache and generate a standardized page table based on the number of PEs and bandwidth of the target hardware. S2. During tensor computation, the DMA component dynamically loads KVCache blocks into the double-buffered weight area, enabling parallelization of data movement and computation. S3. In the decoding phase, select multi-batch collaboration or independent core allocation strategies based on the data scale, and combine mask generation with sparse computing optimization to improve inference efficiency.

8. The compilation optimization method supporting Page Attention according to claim 7, characterized in that: In step S2, the double buffer weight area is divided into a ping area and a pong area. When the tensor component reads the data in the ping area, the DMA component preloads the next batch of KV Cache blocks into the pong area.

9. The compilation optimization method supporting Page Attention according to claim 8, characterized in that: In step S3, the multi-batch collaboration strategy specifically includes: The KV Cache blocks of multiple batches are spliced ​​into the same weight buffer. Multiple computing cores share data and perform parallel calculations. The mask vector only retains the valid data of the current batch.

10. A computer-readable storage medium storing a computer program, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 7 to 9 are implemented.

Citation Information

Cited By

  • Data processing method based on artificial intelligence processor and related product

    CN121614644A

  • Data storage method for large model, medium and computer program product

    CN121832857A