Task scheduling method and device for attention calculation, medium, equipment and product

By employing a task scheduling method based on fine-grained decomposition and asynchronous submission, the high computational complexity and resource consumption issues in attention computation are addressed, thereby improving computational efficiency.

CN121050867AActive Publication Date: 2025-12-02SHANGHAI BIREN TECH CO LTD

Patent Information

Application Number
CN202511596848.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2025-12-02
Estimated Expiration
2045-11-04

AI Technical Summary

Technical Problem

With the development of artificial intelligence technology, the complexity and computation time of attention calculation have increased dramatically, resulting in excessive hardware resource consumption. How to deeply explore the potential of parallelism and optimize the computing process to improve efficiency has become an urgent problem to be solved.

Method used

By finely decomposing the query block loading, attention score block calculation, and intermediate accumulation block update tasks in the attention calculation process, and using multiple task slots in the synchronous channel for asynchronous submission and polling execution, the dependencies between subtasks are managed, and the waiting instruction threshold is dynamically configured to control the execution sequence.

Benefits of technology

It significantly improves the efficiency of attention computing, especially in tasks with the highest computational complexity and the most intensive hardware resource consumption, and achieves performance improvement through fine-grained scheduling optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121050867A_ABST
    Figure CN121050867A_ABST
Patent Text Reader

Abstract

The invention discloses a task scheduling method and device for attention calculation, a medium, equipment and a product. The method comprises the steps that a loading task of a query block is decomposed into N1 loading subtasks, and a calculation task of a jth attention score block is decomposed into N2 calculation subtasks; decomposing an updating task of a middle accumulation block corresponding to the attention output block in the current iteration round into N3 parallel updating sub-tasks; asynchronously submitting all the subtasks to different task slots of N4 synchronization channels; and performing polling operation on the task slot in each synchronization channel to execute the sub-tasks stored in the task slot, and controlling an execution time sequence among the sub-tasks with the dependency relationship in each synchronization channel by configuring a number threshold value of the uncompleted sub-tasks in the waiting instruction. According to the method, the covering capability and the parallel computing capability of hardware can be improved, and the attention computing efficiency is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a task scheduling method, apparatus, computer-readable storage medium, electronic device, and computer program product for attention computation. Background Technology

[0002] With the rapid development of artificial intelligence technology, the attention mechanism has become a core operator in many fields such as natural language processing and computer vision. It achieves dynamic weighted aggregation of input data by calculating the interaction relationships between queries, keys, and values. However, as the model size and the length of the input sequence continue to increase, the complexity of attention computation also increases significantly, leading to a surge in computation time and excessive hardware resource consumption. Therefore, how to deeply explore the parallel potential of attention computation and optimize the computation process to improve computational efficiency is an urgent problem to be solved. Summary of the Invention

[0003] The purpose of this invention is to provide a task scheduling method, apparatus, computer-readable storage medium, electronic device, and computer program product for attention computation. By finely decomposing the loading task of the query block, the calculation task of the attention score block, and the update task of the intermediate accumulation block in the attention computation process, and by utilizing multiple task slots of the synchronous channel to achieve asynchronous submission and polling execution, and by managing the dependencies between subtasks through dynamically configuring threshold parameters for waiting instructions, the hardware's masking ability and parallel computing ability can be improved, thereby improving the efficiency of attention computation.

[0004] A first aspect of the present invention provides a task scheduling method for attention computation, applied to a consumer thread group, wherein the consumer thread group is configured with one or more synchronization channels, and each synchronization channel is configured with multiple task slots; the method includes: The task of loading the query block from shared memory to registers is decomposed into N1 loading subtasks, and the task of computing the j-th attention score block corresponding to the query block is decomposed into N2 computing subtasks. The update task of the intermediate accumulation block corresponding to the attention output block in the current iteration round is decomposed into N3 parallel update subtasks; N1 loading subtasks, N2 computation subtasks, and N3 update subtasks are asynchronously submitted to different task slots of N4 synchronization channels; wherein, subtasks with dependencies are submitted to the same synchronization channel; j≥1, N1≥1, N2≥1, N3≥1, N4≥1; Within each synchronization channel, a polling operation is performed on the task slots to execute the subtasks stored in the task slots. By configuring a threshold for the number of incomplete subtasks in the waiting instructions, the execution sequence of dependent subtasks within each synchronization channel is controlled.

[0005] Optionally, the step of asynchronously submitting the N1 loading subtasks, N2 computation subtasks, and N3 update subtasks to different task slots of the N4 synchronization channels includes: When j=1, N1 loading subtasks and N2 computation subtasks are asynchronously submitted to different task slots in the first synchronization channel group according to their dependencies. When j > 1, N2 computational subtasks are asynchronously submitted to different task slots in the first synchronization channel group; N3 update subtasks are asynchronously submitted to different task slots in the second synchronization channel group; wherein both the first synchronization channel group and the second synchronization channel group consist of at least one of the synchronization channels.

[0006] Optionally, each subtask consists of at least one execution instruction.

[0007] Optionally, the i-th loading subtask is used to load the i-th query sub-block in the query block from shared memory into a register; where 1≤i≤N1; the size of the query sub-block satisfies the first computation size of the matrix multiply-add instruction.

[0008] Optionally, the task of calculating the j-th attention score block corresponding to the query block is decomposed into N2 subtasks, including: The second computational size based on matrix multiplication and addition instructions divides the j-th key transpose block into N1×M key transpose sub-blocks; where M≥1; the size of the key transpose sub-blocks supports the horizontal splicing and expansion of multiple second computational sizes; When j=1, the indexes of the key transpose subblocks are traversed in row-major order. The matrix multiplication and addition of the n1th query subblock and the key transpose subblock in the n1th row and mth column is taken as a computational subtask to form N2 computational subtasks; where 1≤n1≤N1, 1≤m≤M, and N2=N1×M.

[0009] Optionally, the step of decomposing the computation task of the j-th attention score block corresponding to the query block into N2 computation subtasks further includes: When j > 1, the indexes of the key transpose subblocks are traversed in column priority order. The combination of N1 query subblocks and all key transpose subblocks in column m is calculated as a computational subtask to form N2 computational subtasks; where N2 = M.

[0010] Optionally, the first attention score block is calculated through the following steps: When the waiting instruction monitor detects that the first query sub-block has been loaded into the register, the computation subtasks corresponding to all key transpose sub-blocks in the first row are started sequentially. When the waiting instruction monitors that the n2nd query subblock is loaded into the register and the computation subtask corresponding to the key transpose subblock in row n2-1 and column m is completed, the computation subtask corresponding to the key transpose subblock in row n2 and column m is started; where 2≤n2≤N1; When all the computational subtasks corresponding to the key transpose subblocks in column m are completed, the mth fraction subblock of the first attention fraction block is obtained.

[0011] Optionally, the method further includes: When the m-th fractional sub-block of the j-th attention score block is completed by waiting for the instruction to monitor, local statistical calculations are performed on the m-th fractional sub-block.

[0012] Optionally, in the current iteration round, the k-th update subtask is used to multiply the indexed fraction block with the corresponding sub-block in the value block, and accumulate the resulting sub-block product to the corresponding region of the scale-aligned intermediate accumulation block to obtain the updated k-th intermediate accumulation sub-block; where k≥1.

[0013] Optionally, the method further includes: In the next iteration, when the kth update subtask is completed by waiting for the instruction to be monitored, the kth intermediate accumulation sub-block is scaled.

[0014] A second aspect of the present invention provides a task scheduling apparatus for attention computation, applied to a consumer thread group, the consumer thread group being configured with one or more synchronization channels, each synchronization channel being configured with multiple task slots; the apparatus includes: The first decomposition module is used to decompose the loading task of the query block from shared memory to register into N1 loading subtasks, and to decompose the calculation task of the j-th attention score block corresponding to the query block into N2 calculation subtasks. The second decomposition module is used to decompose the update task of the intermediate accumulation block corresponding to the attention output block in the current iteration into N3 parallel update subtasks. The task submission module is used to asynchronously submit N1 loading subtasks, N2 calculation subtasks, and N3 update subtasks to different task slots of the N4 synchronization channels; wherein, subtasks with dependencies are submitted to the same synchronization channel; j≥1, N1≥1, N2≥1, N3≥1, N4≥1; The task execution module is used to poll the task slots in each synchronization channel to execute the subtasks stored in the task slots, and to control the execution sequence of dependent subtasks in each synchronization channel by configuring a threshold for the number of incomplete subtasks in the waiting instructions.

[0015] A third aspect of the present invention provides a computer-readable storage medium comprising a stored computer program; wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the task scheduling method for attention computation as described in any embodiment of the first aspect.

[0016] A fourth aspect of the present invention provides a computer program product including computer instructions, which, when executed by a processor, implement the task scheduling method for attention computation described in any embodiment of the first aspect.

[0017] A fifth aspect of the present invention provides an electronic device including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the task scheduling method for attention computation as described in any embodiment of the first aspect.

[0018] Compared with existing technologies, embodiments of the present invention provide a task scheduling method, apparatus, computer-readable storage medium, electronic device, and computer program product for attention computation. In the attention computation process, embodiments of the present invention perform fine-grained decomposition of the three stages: query block loading, attention score block calculation, and intermediate accumulation block update. Through asynchronous submission of multi-task slots and dependency management based on waiting instructions, the masking capability and parallel computing capability of the hardware are enhanced, making the attention computation pipeline more compact and ultimately significantly improving the computational efficiency of attention. Furthermore, the tasks of calculating attention score blocks and updating intermediate accumulation blocks are the most computationally complex and resource-intensive tasks in the attention computation process. Embodiments of the present invention perform fine-grained scheduling optimization on these tasks, achieving significant performance improvements with relatively small scheduling overhead. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating an embodiment of the task scheduling method for attention computation provided by the present invention; Figure 2 This is a schematic diagram of an embodiment of the block association size in attention calculation provided by the present invention; Figure 3 This is a flowchart illustrating an embodiment of the calculation of attention score blocks provided by the present invention; Figure 4 This is a schematic diagram of another embodiment of the attention score block provided by the present invention.

[0020] Figure 5 This is a schematic diagram of yet another embodiment of the attention score block calculation provided by the present invention; Figure 6 This is a schematic diagram of an embodiment of the intermediate accumulation block update provided by the present invention; Figure 7 This is a schematic diagram of an embodiment of the intermediate accumulation block scale alignment provided by the present invention; Figure 8 This is a schematic diagram of an embodiment of the task scheduling device for attention calculation provided by the present invention; Figure 9 This is a schematic diagram of the structure of an embodiment of the electronic device provided by the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] The artificial intelligence processor involved in this invention can be any one of CPU (Central Processing Unit), GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), NPU (Neural Network Processing Unit), DPU (Deep Learning Processing Unit), APU (Accelerated Processing Unit), and GPGPU (General-Purpose computing on Graphics Processing Unit), depending on its application to a specific product or technology in the embodiments of this invention.

[0023] In this embodiment of the invention, the computational size of the matrix multiply accumulate instruction mma is a×b×K. B Among them, K BThe number of bytes is 32B to 256B; when the matrix data type is FP16, a×b×K B =64×16×16c (c=1, 2,…, 8), meaning the first computational dimension of matrix A is a×K. B =64×16c, the second computational dimension of matrix B is K. B ×b=16c×16. Based on this, in "QK T In multiplication, query blocks The query sub-block in the matrix belongs to matrix A, with a maximum size of 64×128, and is divided by key transpose. If the key transpose sub-block in the matrix belongs to the B matrix, then the maximum size of the key transpose sub-block is 128×16; in "PV" multiplication, the exponential fractional block... The sub-blocks in the matrix A are value-blocked. The sub-blocks in the matrix belong to matrix B; where i≥1; j≥1. In the mma calculation of this embodiment, matrix A is in the thread level register (TLR), and matrix B is in group shared memory (GSM).

[0024] Furthermore, this invention employs the Flash Attention method for attention calculation, decomposing a large matrix into many smaller blocks for computation through a "blocking" approach. In this embodiment, after the producer thread group loads the query block (Q), key block (K), and value block (V) from Global Memory (GLM) to GSM in batches, the consumer thread group performs attention calculation through the following main operations / processes: (1) In the outer loop of Q, the query is divided into blocks. Loaded into the TLR for iterative computation within the KV inner loop to generate the i-th attention output block. ); (2) In the iterative calculation of the KV inner loop, the query blocks in the TLR are processed by the specified matrix multiplication and addition instruction (mma). Key block in GSM Conduct "QK" T Multiplication (i.e.) ), to obtain the corresponding attention score blocks ; (3) Obtain attention score blocks The current maximum value tensor And combined with the historical cumulative maximum tensor Calculate the current cumulative maximum value tensor ; (4) Obtain the update correction factor Modifier= Used for scale alignment; (5) Segmenting attention scores Perform exponential operations (i.e.) ), to obtain indexed fractional blocks Where sf is the scaling factor; (6) Obtain the indexed score blocks The current exponent and tensor And combining historical accumulation and tensors Calculate the current accumulated sum tensor ; (7) Segmenting attention output Corresponding intermediate accumulation blocks Perform scale alignment operation, i.e. = ; (8) Blocking of indexed fractions Sum value block Perform "PV" multiplication (i.e.) Then, the product results are accumulated into the scale-aligned intermediate accumulation blocks (i.e., ...). ), to obtain the intermediate cumulative blocks after iterative updates ,Right now Scale-aligned intermediate cumulative blocks = ; (9) After the inner loop iteration calculation is completed, calculate the current accumulated sum tensor obtained from the last iteration. Accumulate blocks in the middle Perform a rescale operation to obtain the attention output blocks. .

[0025] The above (2) to (8) are located in the iterative calculation of the KV inner loop, while (1) and (9) are located in the Q outer loop.

[0026] The embodiments of the present invention perform fine-grained scheduling optimization on the above attention calculation process, and propose a task scheduling method, apparatus, computer-readable storage medium, electronic device and computer program product for attention calculation.

[0027] See Figure 1 This is a flowchart illustrating an embodiment of the task scheduling method for attention computation provided by the present invention.

[0028] A first aspect of the present invention provides a task scheduling method for attention computation, applied to a consumer thread group, wherein the consumer thread group is configured with one or more synchronization channels, and each synchronization channel is configured with multiple task slots; the method includes steps S1 to S4, as follows: Step S1: Decompose the task of loading the query block from shared memory to register into N1 loading subtasks, and decompose the task of calculating the j-th attention score block corresponding to the query block into N2 calculation subtasks. Step S2: Decompose the update task of the intermediate accumulation block corresponding to the attention output block in the current iteration into N3 parallel update subtasks; Step S3: Asynchronously submit N1 loading subtasks, N2 computation subtasks, and N3 update subtasks to different task slots of the N4 synchronization channels; wherein, subtasks with dependencies are submitted to the same synchronization channel; j≥1, N1≥1, N2≥1, N3≥1, N4≥1; Step S4: Poll the task slots in each synchronization channel to execute the subtasks stored in the task slots, and control the execution sequence of dependent subtasks in each synchronization channel by configuring a threshold for the number of incomplete subtasks in the waiting instructions.

[0029] For example, for a certain artificial intelligence chip, the consumer thread group is configured with 3 synchronization channels (SyncChannel), namely sc0, sc1 and sc2; each synchronization channel is configured with 8 task slots (slot0~slot7).

[0030] The embodiments of the present invention, through the aforementioned hardware conditions, perform query segmentation. The loading phase (corresponding to operation (1) above), attention score block The calculation phase (corresponding to operation (2) above) involves "QK T Multiplication), intermediate accumulation blocks The iterative update phase (corresponding to the operation (8) above, involving “PV” multiplication) is optimized and executed with fine granularity.

[0031] During the query chunk loading phase, the query chunks are... The loading task from GSM to TLR is broken down into N1 loading subtasks. For example, a query block of size 64×256 is divided into subtasks. (Data type is FP16) Logically divided into 2 query sub-blocks (size is 64×128), then for the query sub-blocks The loading task is then broken down into N1=2 loading subtasks, each responsible for loading the corresponding query subblock into the register, so as to further implement "QK" through the mma instruction. T "multiplication.

[0032] Attention score segmentation During the computation phase, key transpose-based block partitioning is performed. Granularity of division (e.g., dividing) The logic is divided into N1×M key transpose sub-blocks, and the attention score is divided into blocks. computational tasks ( The matrix is ​​decomposed into one or more query sub-blocks and corresponding key transpose sub-blocks.

[0033] For example, transpose a 256×32 key into blocks. The logic is divided into N1×M=2×2 key transpose sub-blocks (size 128×16), corresponding to attention score blocks. There are two ways to decompose the computational task: ① Decompose it into 4 computational subtasks, each of which is a matrix multiplication and addition calculation of one query subblock and one key transpose subblock (here, one computational subtask corresponds to one mma instruction). This decomposition method can effectively mask the gap between data loading and computation when processing the first key transpose block; ② Decompose it into 2 computational subtasks, each of which is a combined calculation of 2 query subblocks and all key transpose subblocks in a single column (here, one computational subtask corresponds to 2 mma instructions). This decomposition method can directly calculate the attention score block. The corresponding column is divided into fractional sub-blocks, which can reduce the overhead of task slots. Both of the above decomposition methods are applicable to all key transpose blocks, but for the case of j=1, the first decomposition method is preferred.

[0034] Accumulate blocks in the middle During the iterative update phase, the update task is... Scale-aligned intermediate accumulation blocks ( ), divide the intermediate accumulation blocks The logic is divided into N3 intermediate accumulation sub-blocks, thus dividing the intermediate accumulation blocks. The update task is decomposed into N3 parallel update subtasks, each of which is used to update the corresponding intermediate accumulator subblock.

[0035] In step S4, N1 loading subtasks, N2 computation subtasks, and N3 update subtasks are asynchronously submitted to different task slots in N4 synchronization channels. This ensures that each subtask is stored in a separate task slot, and subtasks with dependencies are submitted to the same synchronization channel to avoid computational delays caused by cross-channel synchronization. In this embodiment of the invention, there is a dependency between the loading subtask and the computation subtask corresponding to the first attention score block (computation requires waiting for the corresponding query subblock to finish loading); there are no dependencies between the loading subtasks; similarly, there are no dependencies between the update subtasks.

[0036] In step S5, a round-robin operation is performed on the task slots in each synchronization channel, and the subtasks in each task slot are started sequentially (e.g., in sc0, the subtasks are executed in a loop from slot0 to slot7), ensuring that the subtasks of all task slots in the synchronization channel are scheduled.

[0037] The wait_group(scX,n_pending) (i.e., the waitle instruction) is used to arrange the execution timing of dependent subtasks; where scX is the synchronization channel numbered X; n_pending is the threshold for the number of unfinished subtasks, ranging from 0 to the total number of slots in the synchronization channel minus 1. For example, if sc0 has 8 slots, then the range of n_pending is 0 to 7.

[0038] The `waitle` directive `wait_group(scX, n_pending)` causes the current thread to wait until the number of pending tasks in the synchronization channel `scX` equals the preset `n_pending`, at which point it is unblocked and continues executing subsequent instructions. For example, in a scenario where `j=1`, configuring the `waitle` directive ensures that computational subtasks wait for the slot containing the preceding subtask to finish loading before execution, avoiding computational errors caused by incomplete data loading.

[0039] As can be seen from the above, the embodiments of the present invention decompose the attention calculation process into three stages: query block loading, attention score block calculation, and intermediate accumulation block update. Through asynchronous submission of multi-task slots and dependency management based on waiting instructions, the hardware characteristics of multi-task slots in the synchronous channel can be fully utilized, enhancing the hardware's masking and parallel computing capabilities. This results in a more compact attention calculation pipeline, ultimately significantly improving the computational efficiency of attention. Furthermore, the attention score block calculation task (involving "QK")... TThe tasks of updating the intermediate accumulation blocks (involving "PV" multiplication) and "multiplication" are the most computationally complex and hardware resource-intensive tasks in the attention computation process. The embodiments of the present invention optimize them with fine-grained scheduling, which can achieve significant performance improvement with a small scheduling overhead.

[0040] In an optional embodiment, the asynchronous submission of N1 loading subtasks, N2 computation subtasks, and N3 update subtasks to different task slots of the N4 synchronization channels includes: When j=1, N1 loading subtasks and N2 computation subtasks are asynchronously submitted to different task slots in the first synchronization channel group according to their dependencies. When j > 1, N2 computational subtasks are asynchronously submitted to different task slots in the first synchronization channel group; N3 update subtasks are asynchronously submitted to different task slots in the second synchronization channel group; wherein both the first synchronization channel group and the second synchronization channel group consist of at least one of the synchronization channels.

[0041] It should be noted that the first synchronization channel group is used to carry the loading subtask (only when j=1) and calculation subtask directly related to attention score calculation; the second synchronization channel group is used to carry the update subtask related to intermediate accumulation block iterative update; the number of channels contained in the two synchronization channel groups can be dynamically configured according to the task size.

[0042] In the processing stage of the first key transpose block (i.e., j=1), since the computation subtask has a strong dependency on the loading subtask (it needs to wait for the query subblock to be loaded into the register before the corresponding mma computation can be started), the loading subtask needs to occupy the preceding task slot of the synchronization channel to ensure that the computation subtask can perceive its execution status through the waittle instruction.

[0043] In the processing stage of the second and subsequent key transpose blocks (i.e., when j > 1), all query sub-blocks in the query block have been loaded into the registers, eliminating the need to repeatedly execute the loading subtasks. Therefore, only N2 computation subtasks need to be asynchronously submitted to different task slots in the first synchronization channel group. In other words, when j > 1, there is no need to reserve slots for loading subtasks, and computation subtasks can be submitted continuously starting from the initial slot in the first synchronization channel group.

[0044] Regardless of whether j=1 or j>1, all N3 update subtasks must be submitted asynchronously to different task slots in the second synchronization channel group. When the number of update subtasks is large (e.g., N3=16), and the 8 slots of a single synchronization channel are insufficient, the second synchronization channel group can be expanded into 2 synchronization channels (e.g., sc1 and sc2), with each synchronization channel carrying 8 update subtasks.

[0045] In an optional embodiment, the i-th loading subtask is used to load the i-th query sub-block in the query block from shared memory into a register; where 1≤i≤N1; the size of the query sub-block satisfies the first computation size of the matrix multiply-add instruction.

[0046] For example, such as Figure 2 The diagram shown is a schematic representation of an embodiment of the block association dimensions in attention calculation provided by the present invention. Figure 2 In this context, Q chunk, K chunk, and V chunk represent the query chunk, key chunk, and value chunk, respectively; S represents the attention score chunk calculated from the query chunk and key chunk; P represents the exponential score chunk of S after softmax; and O chunk is the attention output chunk corresponding to the query chunk.

[0047] As mentioned earlier, when the data type is FP16, the first computation size of the matrix multiply-add instruction (mma) is defined as 64×16c (c=1, 2,…, 8). Figure 2 The size of the Q chunk is 64×256, therefore the Q chunk can be logically divided into N1=2 query sub-blocks, each with a size of 64×128 (where c=8); for example Figure 3 The diagram shown is a schematic representation of an embodiment of the attention score block calculation provided by the present invention. Figure 3 In this context, "Half Q" represents a query subblock (64×128 in size), and "Half K" represents a key transpose subblock (128×16 in size).

[0048] In an optional embodiment, the task of decomposing the calculation task of the j-th attention score block corresponding to the query block into N2 calculation subtasks includes: The second computational size based on matrix multiplication and addition instructions divides the j-th key transpose block into N1×M key transpose sub-blocks; where M≥1; the size of the key transpose sub-blocks supports the horizontal splicing and expansion of multiple second computational sizes; When j=1, the indexes of the key transpose subblocks are traversed in row-major order. The matrix multiplication and addition of the n1th query subblock and the key transpose subblock in the n1th row and mth column is taken as a computational subtask to form N2 computational subtasks; where 1≤n1≤N1, 1≤m≤M, and N2=N1×M.

[0049] like Figure 3As shown, when the data type is FP16, the second computation size of the mma instruction is defined as 16c×16. Combined with the query sub-block size of 64×128, and the number of task slots in a single synchronization channel meeting the computation scheduling requirements of the current attention score block, the key transpose block (size 256×32) is logically divided into N1×M=2×2 key transpose sub-blocks. Each key transpose sub-block has a size of 128×16 (the second computation size when c=8), and no horizontal expansion of the second computation size is required.

[0050] In the processing stage of the first key-transposed block (j=1), the indices of the key-transposed sub-blocks are traversed in row-major order. The matrix multiplication and addition calculation of the n1-th query sub-block and the key-transposed sub-block in the n1-th row and m-th column is treated as an independent computational subtask, forming N2=N1×M computational subtasks. At this time, each computational subtask corresponds to one mma instruction, which can ensure that, in the case of j=1, the gap between data loading and computation is fully masked by fine-grained task decomposition.

[0051] by Figure 3 Taking the scenario where N1=2 and M=2 as an example, the matrix multiplication and addition calculation of the first query sub-block (the previous "Half Q" / fronthalfQ) and the key transposed sub-block in the first row and first column is regarded as the first calculation sub-task, and the matrix multiplication and addition calculation of the key transposed sub-block in the first row and second column is regarded as the second calculation sub-task; the mma calculation of the second query sub-block (the next "Half Q" / back halfQ) and the key transposed sub-block in the second row and first column is regarded as the third calculation sub-task, and the mma calculation of the key transposed sub-block in the second row and second column is regarded as the fourth calculation sub-task.

[0052] The consumer thread group implements this through instruction example 1. Figure 3 The calculation of the first attention score block is shown below. Of course, the task decomposition method in Instruction Example 1 can also be used to calculate the second and subsequent attention score blocks. Specific Instruction Example 1 is shown below: Loop{ / KV internal circulation / ldmatrix front halfQ(2instr) / / The first load subtask, loads the first query subblock (fronthalfQ) from GSM into TLR. This load subtask contains 2 ldmatrix instructions; commit sc0 / / Submit the first loading subtask to slot0 of sc0; ldmatrix back halfQ(2instr) / / The second loading subtask loads the second query subblock (backhalfQ) from GSM into TLR, which also contains 2 ldmatrix instructions; commit sc0 / / Submit the second loading subtask to slot1 of sc0; waitle sc0, 1 / / Wait until the number of unfinished subtasks in sc0 equals 1, then unblock and continue execution. At this point, the first query subblock has been loaded into the TLR. regS1=front mma(halfQ,halfK) / / The first computational subtask performs mma computation on the first query subblock and the key transpose subblock (halfK) in the first row and first column, and stores the result in register regS1; commit sc0 / / Submit the first computational subtask to slot2 of sc0; regS2 = front mma(halfQ, halfK) / / The second computation subtask, performs mma computation on the first query subblock and the key transpose subblock in the first row and second column, and stores the result in regS2; commit sc0 / / Submit the second computational subtask to slot 3 of sc0; waitle sc0, 1 / / Wait until the number of unfinished subtasks in sc0 equals 1, then execute the 3rd computation subtask, ensuring that the result data provided after the 1st computation subtask is completed is used for subsequent accumulation operations; regS1 = mma(halfQ, halfK,regS1) / / The third computational subtask, performs mma computation on the second query sub-block and the key transpose sub-block in the second row and first column, and adds the result to regS1 to obtain the first score sub-block of the first attention score block; commit sc0 / / Submit the third computational subtask to slot 4 of sc0; waitle sc0, 1 / / Wait until the number of unfinished subtasks in sc0 equals 1, then execute the 4th calculation subtask, ensuring that the result data provided after the 2nd calculation subtask is completed is used for subsequent accumulation operations; regS2 = mma(halfQ, halfK,regS2) / / The 4th computational subtask, performs mma computation on the 2nd query sub-block and the key transpose sub-block in the 2nd row and 2nd column, and adds the result to regS2 to obtain the 2nd score sub-block of the 1st attention score block; commit sc0 / / Submit the 4th computational subtask to slot 5 of sc0; waitle sc0, 1 / / Wait until the number of unfinished subtasks in sc0 equals 1, then the first fractional sub-block is completed and can be used to perform local statistical calculations, such as finding the maximum value in the row direction (row_max instruction). row_max(regS1) waitle sc0, 0 / / Wait until the number of unfinished subtasks in sc0 equals 0, then the second fractional sub-block is completed and local statistical calculations can be performed on it; row_max(regS2) … } Based on the above embodiments, the first attention score block is calculated through the following steps: When the waiting instruction monitor detects that the first query sub-block has been loaded into the register, the computation subtasks corresponding to all key transpose sub-blocks in the first row are started sequentially. When the waiting instruction monitors that the n2nd query subblock is loaded into the register and the computation subtask corresponding to the key transpose subblock in row n2-1 and column m is completed, the computation subtask corresponding to the key transpose subblock in row n2 and column m is started; where 2≤n2≤N1; When all the computational subtasks corresponding to the key transpose subblocks in column m are completed, the mth fraction subblock of the first attention fraction block is obtained.

[0053] Specifically, with Figure 3 For example, in the corresponding instruction example 1, when the waiter instruction detects that the first query subblock (fronthalfQ) is loaded from GSM to TLR, it immediately starts the computation subtasks corresponding to all key transpose subblocks in the first row (i.e., the first computation subtask and the second computation subtask).

[0054] When the waittle instruction detects that the second query sub-block (back halfQ) has been loaded and the first computation subtask has been completed, it immediately starts the third computation subtask. Similarly, when the waittle instruction detects that the second query sub-block (back halfQ) has been loaded and the second computation subtask has been completed, it immediately starts the fourth computation subtask.

[0055] Figure 3After the computational subtasks corresponding to the key transpose subblocks in the first row and first column and the second row and first column (i.e., the first and third computational subtasks) are completed, the accumulated result stored in register regS1 is the first fractional subblock. Similarly, after the computational subtasks corresponding to the key transpose subblocks in the first row and second column and the second row and second column (i.e., the second and fourth computational subtasks) are completed, the accumulated result stored in register regS2 is the second fractional subblock.

[0056] In an optional embodiment, when j > 1, the indexes of the key transposed sub-blocks are traversed in column priority order, and the combination of N1 query sub-blocks and all key transposed sub-blocks in column m is calculated as a computational subtask to form N2 computational subtasks; where N2 = M.

[0057] like Figure 4 The diagram shown is a schematic of another embodiment of the attention score block provided by the present invention. In the processing stage of the second and subsequent key transpose blocks (j>1), the query sub-block has been stored in the TLR through the loading subtask of the j=1 stage and does not need to be loaded again. Therefore, there is no need to mask the loading and calculation. The mma instructions can be integrated at the column level to obtain the corresponding calculation subtask.

[0058] by Figure 4 Taking the scenario where N1=2 and M=2 as an example, the matrix multiplication and addition calculations of the first query sub-block (the previous "Half Q") and the key transpose sub-block in the first row and first column, and the matrix multiplication and addition calculations of the second query sub-block (the next "Half Q") and the key transpose sub-block in the second row and first column, are collectively considered as the first computational subtask; the matrix multiplication and addition calculations of the first query sub-block and the key transpose sub-block in the first row and second column, and the matrix multiplication and addition calculations of the second query sub-block and the key transpose sub-block in the second row and second column, are collectively considered as the second computational subtask. In this case, each computational subtask corresponds to two mma instructions.

[0059] The consumer thread group implements this through instruction example 2. Figure 4 The calculation of the second and subsequent attention score blocks. Example instruction 2 is shown below: Loop{ / KV internal circulation / waitle sc0, 1 / / Release the block when the number of unfinished subtasks in sc0 equals 1; regS1=front mma(halfQ,halfK) / / An mma instruction for the first computation subtask, which performs mma computation on the first query subblock and the key transpose subblock in the first row and first column, and stores the result in regS1; regS1 = mma(halfQ,halfK,regS1) / / Another mma instruction for the first computational subtask, which performs mma computation on the second query sub-block and the key transpose sub-block in the second row and first column, and accumulates the result into regS1 to obtain the first score sub-block of the attention score block; commit sc0 / / Submit the first computational subtask to slot0 of sc0; regS2 =front mma(halfQ,halfK) / / An mma instruction for the second computation subtask, which performs mma computation on the first query subblock and the key transpose subblock in the first row and second column, and stores the result in regS2; regS2 = mma(halfQ,halfK,regS2) / / Another mma instruction for the second computational subtask, which performs mma computation on the second query sub-block and the key transpose sub-block in the second row and second column, and adds the result to regS2 to obtain the second score sub-block of the attention score block; commit sc0 / / Submit the second computational subtask to slot1 of sc0; waitle sc0, 1 / / Wait until the number of unfinished subtasks in sc0 equals 1, then the first fractional sub-block is completed and can be used to perform local statistical calculations, such as finding the maximum value in the row direction (row_max instruction). row_max(regS1) waitle sc0, 0 / / Wait until the number of unfinished subtasks in sc0 equals 0, then the second fractional sub-block is completed and local statistical calculations can be performed on it; row_max(regS2) … } As can be seen from instruction example 1 and instruction example 2, each subtask consists of at least one execution instruction.

[0060] In an optional embodiment, the method further includes: When the m-th fractional sub-block of the j-th attention score block is completed by waiting for the instruction to monitor, local statistical calculations are performed on the m-th fractional sub-block.

[0061] As shown in Instruction Examples 1 and 2, once the waittle instruction confirms that the m-th fractional sub-block has been calculated, the register storing the m-th fractional sub-block (e.g., regS1 corresponds to the 1st fractional sub-block, regS2 corresponds to the 2nd fractional sub-block) is accessed directly. The row_max(regSm) instruction is called to calculate the maximum value in the row direction for the fractional sub-block in register regSm, thus obtaining the local maximum value. After calculating the local maximum values ​​of all fractional sub-blocks, the global maximum value of the current attention fractional block is obtained.

[0062] It is worth noting that when the number of task slots in a single synchronization channel cannot meet the computational scheduling requirements of the current attention score block, the size of the key transpose sub-block in this embodiment of the invention supports the horizontal splicing and expansion of multiple second computational sizes, thereby reducing the consumption of task slots.

[0063] like Figure 5 The diagram shown is a schematic of another embodiment of the attention score block provided by the present invention. The size of the query sub-block is 64×96. If no splicing expansion is performed, the key transpose block (size 192×64) is logically divided into N1×M=2×4 key transpose sub-blocks, and the size of each key transpose sub-block is 96×16 (the second calculation size when c=6).

[0064] according to Figure 3 As shown in the corresponding instruction example 1, with N1×M=2×4 key transpose sub-blocks, 6 task slots of sc0 are required (2 loading sub-tasks + 4 calculation sub-tasks). Therefore, for... Figure 5 With the configuration of N1×M=2×4 key transpose sub-blocks, the required number of task slots increases to 10 (2 loading sub-tasks + 8 calculation sub-tasks), exceeding the total number of slots in sc0 (8). At this point, the 10 sub-tasks can be split into 2 batches for calculation (i.e., wait for the previous batch of sub-tasks to be fully or partially completed before submitting the next batch of sub-tasks), or, if resources are sufficient, an additional synchronization channel can be used to complete the calculation.

[0065] Of course, it is also possible to decompose the task in the first attention score block computation stage, and then... Figure 5 The key transpose block logic is divided into N1×M=2×2 key transpose sub-blocks to ensure that the number of task slots occupied (6) does not exceed the total number of slots. Each key transpose sub-block is 96×32 pixels, obtained by horizontally concatenating two second calculation dimensions. Therefore, two mma instructions are needed to complete the matrix multiplication and addition calculations for one query sub-block and one key transpose sub-block. In other words, in the calculation... Figure 5 The first attention score block requires four computation tasks, and each computation task contains two mma instructions.

[0066] In an optional embodiment, in the current iteration round, the k-th update subtask is used to multiply the indexed fraction block with the corresponding sub-block in the value block, and accumulate the resulting sub-block product to the corresponding region of the scale-aligned intermediate accumulation block to obtain the updated k-th intermediate accumulation sub-block; where k≥1.

[0067] Furthermore, the method also includes: In the next iteration, when the kth update subtask is completed by waiting for the instruction to be monitored, the kth intermediate accumulation sub-block is scaled.

[0068] like Figure 6 The diagram shown is a schematic of an embodiment of the intermediate accumulation block update provided by the present invention. The size of the exponential fraction block ("P") is 64×32, which directly matches the first calculation size of the mma instruction and does not need to be divided; the size of the value block is 32×256, which needs to be logically divided into 16 sub-blocks according to the second calculation size of mma. Each sub-block (size 32×16) corresponds to an update subtask. Therefore, the update task of the intermediate accumulation block is decomposed into 16 update subtasks that can be executed in parallel.

[0069] Since a single synchronization channel has 8 slots, in the current iteration round (denoted as the j-th round), the update operation of the intermediate accumulation block (corresponding to the operation (8) mentioned above, involving "PV" multiplication) requires 2 synchronization channels (sc1 and sc2); sc1's 8 slots are used to carry the 8 update subtasks corresponding to the first half of the value block HalfV (containing 8 sub-blocks); sc2's 8 slots are used to carry the 8 update subtasks corresponding to the second half of the value block HalfV (containing 8 sub-blocks); each slot stores only one update subtask. After all the update subtasks of sc1 are executed, the updated first half of the intermediate accumulation block can be obtained in register regO1; similarly, after all the update subtasks of sc2 are executed, the updated second half of the intermediate accumulation block can be obtained in register regO2.

[0070] like Figure 7 The diagram shown is a schematic representation of an embodiment of the intermediate accumulation block scale alignment provided by the present invention. Before iteratively updating the intermediate accumulation blocks in the next iteration (denoted as the (j+1)th iteration), the intermediate accumulation blocks of the jth iteration need to be scale-aligned. Figure 7 Taking sc1 as an example, each yellow block represents an update subtask of sc1 in the j-th round (occupying 1 slot); each blue block represents the scale alignment operation of the k-th intermediate cumulative sub-block obtained in the j-th round in the (j+1)-th round (corresponding to the operation (7) above).

[0071] Specifically, in the (j+1)th round, when the waittle instruction monitors and detects that the kth update subtask of sc1 has been completed, it indicates that the kth intermediate accumulator subblock in the jth round has been updated. Immediately, the update correction factor Modifier in the (j+1)th round is used to scale it so as to further execute the kth update subtask in the (j+1)th round.

[0072] In summary, the embodiments of the present invention achieve a temporal overlap between the update subtask and the alignment operation by asynchronously submitting the update subtask and waiting for the scale alignment to be triggered. For example... Figure 7 As shown, the update task submitted to the synchronization channel sc1 in the j-th round can partially overlap with the scale alignment task of sc1 in the (j+1)-th round in time, without having to wait for all update subtasks in the j-th round to be completed before performing subsequent operations.

[0073] It's worth noting that when hardware synchronization channel resources are limited—for example, if these resources are occupied by other tasks, or if the hardware itself has too few synchronization channels to meet the independent grouping requirements—the first and second synchronization channel groups are allowed to share the same synchronization channel. This is because the update subtask starts after the loading / computation subtask has finished (the two are not executed at the same time). At this point, the scale alignment operation of the intermediate accumulation block in round j+1 cannot be refined and must wait for all computation subtasks to complete. Furthermore, if the intermediate accumulation matrix in round j is updated in batches, the scale alignment operation of the intermediate accumulation block in round j+1 also needs to wait for all computation subtasks to complete.

[0074] See Figure 8 This is a schematic diagram of an embodiment of the task scheduling device for attention calculation provided by the present invention.

[0075] A second aspect of the present invention provides a task scheduling apparatus for attention computation, applied to a consumer thread group, the consumer thread group being configured with one or more synchronization channels, each synchronization channel being configured with multiple task slots; the apparatus includes: The first decomposition module 11 is used to decompose the loading task of the query block from shared memory to register into N1 loading subtasks, and to decompose the calculation task of the j-th attention score block corresponding to the query block into N2 calculation subtasks. The second decomposition module 12 is used to decompose the update task of the intermediate accumulation block corresponding to the attention output block in the current iteration round into N3 parallel update subtasks. The task submission module 13 is used to asynchronously submit N1 loading subtasks, N2 calculation subtasks, and N3 update subtasks to different task slots of the N4 synchronization channels; wherein, subtasks with dependencies are submitted to the same synchronization channel; j≥1, N1≥1, N2≥1, N3≥1, N4≥1; The task execution module 14 is used to poll the task slots in each synchronization channel to execute the subtasks stored in the task slots, and to control the execution sequence of dependent subtasks in each synchronization channel by configuring a threshold for the number of incomplete subtasks in the waiting instructions.

[0076] It should be noted that the task scheduling device for attention computing provided in the second aspect embodiment of the present invention can implement all the processes of the task scheduling method for attention computing described in any of the first aspect embodiments. The functions and technical effects of each module and unit in the device are the same as the functions and technical effects of the task scheduling method for attention computing described in any of the first aspect embodiments, and will not be repeated here.

[0077] A third aspect of the present invention provides a computer-readable storage medium comprising a stored computer program; wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the task scheduling method for attention computation as described in any embodiment of the first aspect.

[0078] A fourth aspect of the present invention provides a computer program product including computer instructions, which, when executed by a processor, implement the task scheduling method for attention computation described in any embodiment of the first aspect.

[0079] See Figure 9 This is a schematic diagram of an embodiment of the electronic device provided by the present invention.

[0080] A fifth aspect of the present invention provides an electronic device including a processor 21, a memory 22, and a computer program stored in the memory 22 and configured to be executed by the processor 21, wherein the processor 21, when executing the computer program, implements the task scheduling method for attention computation as described in any embodiment of the first aspect.

[0081] Preferably, the computer program can be divided into one or more modules / units (such as computer program one, computer program two, ...), and the one or more modules / units are stored in the memory 22 and executed by the processor 21 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0082] The processor 21 can be any one of a CPU (Central Processing Unit), GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), NPU (Neural Network Processing Unit), DPU (Deep Learning Processing Unit), APU (Accelerated Processing Unit), and GPGPU (General-Purpose Computing on Graphics Processing Unit). The processor 21 is the control center of the electronic device, connecting various parts of the electronic device via various interfaces and lines.

[0083] The memory 22 mainly includes a program storage area and a data storage area. The program storage area can store the operating system, applications required for at least one function, etc., and the data storage area can store related data, etc. In addition, the memory 22 can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, and a flash card, etc., or the memory 22 can also be other volatile solid-state storage devices.

[0084] It should be noted that the aforementioned electronic devices may include, but are not limited to, processors and memory, as will be understood by those skilled in the art. Figure 9 The structural block diagram shown is merely a structural example of the above-described electronic device and does not constitute a limitation on the structure of the above-described electronic device. The above-described electronic device may include more or fewer components than shown, or combine certain components, or different components.

[0085] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A task scheduling method for attention computation, characterized in that, Applied to a consumer thread group, the consumer thread group being configured with one or more synchronization channels, each synchronization channel being configured with multiple task slots; the method includes: The task of loading the query block from shared memory to registers is decomposed into N1 loading subtasks, and the task of computing the j-th attention score block corresponding to the query block is decomposed into N2 computing subtasks. The update task of the intermediate accumulation block corresponding to the attention output block in the current iteration round is decomposed into N3 parallel update subtasks; N1 loading subtasks, N2 computation subtasks, and N3 update subtasks are asynchronously submitted to different task slots of N4 synchronization channels; wherein, subtasks with dependencies are submitted to the same synchronization channel; j≥1, N1≥1, N2≥1, N3≥1, N4≥1; Within each synchronization channel, a polling operation is performed on the task slots to execute the subtasks stored in the task slots. By configuring a threshold for the number of incomplete subtasks in the waiting instructions, the execution sequence of dependent subtasks within each synchronization channel is controlled.

2. The task scheduling method for attention calculation as described in claim 1, characterized in that, The step of asynchronously submitting N1 loading subtasks, N2 computation subtasks, and N3 update subtasks to different task slots of the N4 synchronization channels includes: When j=1, N1 loading subtasks and N2 computation subtasks are asynchronously submitted to different task slots in the first synchronization channel group according to their dependencies. When j > 1, N2 computational subtasks are asynchronously submitted to different task slots in the first synchronization channel group; N3 update subtasks are asynchronously submitted to different task slots in the second synchronization channel group; wherein both the first synchronization channel group and the second synchronization channel group consist of at least one of the synchronization channels.

3. The task scheduling method for attention calculation as described in claim 1, characterized in that, Each subtask consists of at least one execution instruction.

4. The task scheduling method for attention calculation as described in claim 1, characterized in that, The i-th loading subtask is used to load the i-th query sub-block from shared memory into a register; where 1≤i≤N1; the size of the query sub-block satisfies the first computation size of the matrix multiply-add instruction.

5. The task scheduling method for attention calculation as described in claim 4, characterized in that, The task of decomposing the calculation task of the j-th attention score block corresponding to the query block into N2 calculation subtasks includes: The second computational size based on matrix multiplication and addition instructions divides the j-th key transpose block into N1×M key transpose sub-blocks; where M≥1; the size of the key transpose sub-blocks supports the horizontal splicing and expansion of multiple second computational sizes; When j=1, the indexes of the key transpose subblocks are traversed in row-major order. The matrix multiplication and addition of the n1th query subblock and the key transpose subblock in the n1th row and mth column is taken as a computational subtask to form N2 computational subtasks; where 1≤n1≤N1, 1≤m≤M, and N2=N1×M.

6. The task scheduling method for attention computation as described in claim 5, characterized in that, The step of decomposing the computation task of the j-th attention score block corresponding to the query block into N2 computation subtasks further includes: When j > 1, the indexes of the key transpose subblocks are traversed in column priority order. The combination of N1 query subblocks and all key transpose subblocks in column m is calculated as a computational subtask to form N2 computational subtasks; where N2 = M.

7. The task scheduling method for attention computation as described in claim 5, characterized in that, The first attention score block is calculated using the following steps: When the waiting instruction monitor detects that the first query sub-block has been loaded into the register, the computation subtasks corresponding to all key transpose sub-blocks in the first row are started sequentially. When the waiting instruction monitors that the n2nd query subblock is loaded into the register and the computation subtask corresponding to the key transpose subblock in row n2-1 and column m is completed, the computation subtask corresponding to the key transpose subblock in row n2 and column m is started; where 2≤n2≤N1; When all the computational subtasks corresponding to the key transpose subblocks in column m are completed, the mth fraction subblock of the first attention fraction block is obtained.

8. The task scheduling method for attention computation as described in claim 5 or 6, characterized in that, The method further includes: When the m-th fractional sub-block of the j-th attention score block is completed by waiting for the instruction to monitor, local statistical calculations are performed on the m-th fractional sub-block.

9. The task scheduling method for attention computation as described in claim 1, characterized in that, In the current iteration round, the k-th update subtask is used to multiply the indexed fraction block with the corresponding sub-block in the value block, and accumulate the resulting sub-block product to the corresponding region of the scale-aligned intermediate accumulation block to obtain the updated k-th intermediate accumulation sub-block; where k≥1.

10. The task scheduling method for attention computation as described in claim 9, characterized in that, The method further includes: In the next iteration, when the kth update subtask is completed by waiting for the instruction to be monitored, the kth intermediate accumulation sub-block is scaled.

11. A task scheduling device for attention calculation, characterized in that, An apparatus applied to a consumer thread group, the consumer thread group being configured with one or more synchronization channels, each of the synchronization channels being configured with multiple task slots; the apparatus includes: The first decomposition module is used to decompose the loading task of the query block from shared memory to register into N1 loading subtasks, and to decompose the calculation task of the j-th attention score block corresponding to the query block into N2 calculation subtasks. The second decomposition module is used to decompose the update task of the intermediate accumulation block corresponding to the attention output block in the current iteration into N3 parallel update subtasks. The task submission module is used to asynchronously submit N1 loading subtasks, N2 calculation subtasks, and N3 update subtasks to different task slots of the N4 synchronization channels; wherein, subtasks with dependencies are submitted to the same synchronization channel; j≥1, N1≥1, N2≥1, N3≥1, N4≥1; The task execution module is used to poll the task slots in each synchronization channel to execute the subtasks stored in the task slots, and to control the execution sequence of dependent subtasks in each synchronization channel by configuring a threshold for the number of incomplete subtasks in the waiting instructions.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program; wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium resides to perform a task scheduling method for attention computation as described in any one of claims 1 to 10.

13. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the task scheduling method for attention computation as described in any one of claims 1 to 10.

14. An electronic device, characterized in that, The system includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the task scheduling method for attention computation as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Attention operation processing method and device

    CN118585249A

  • Self-adaptive task scheduling execution unit management method and system

    CN119376903A

  • Distributed training method, device and equipment for large-scale model

    CN120409554A

Cited By

  • Distributed attention calculation system

    CN122019113A

  • A distributed attention computing system

    CN122019113B