Attention mechanism operation method, electronic equipment and storage medium
Patent Information
- Application Number
- CN202510378793.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-13
Smart Images

Figure CN120144182A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence chips, and particularly relates to a method for operating an attention mechanism, an electronic device, and a storage medium. Background Art
[0002] In a scenario of performing calculations based on an artificial intelligence chip, multiple warps are usually paralleled in an execution unit (EU). The parallelism of multiple warps can mutually cover the idle time in instruction execution, thereby improving the overall execution efficiency of the execution unit.
[0003] However, limited by computing resources, the execution unit can only issue a few instructions at the same time. This causes the instructions to be issued successively based on a preset rule when multiple warps all have the condition for instruction issuance, resulting in out-of-order execution of multiple warps in the execution unit.
[0004] Specifically, in a scenario of operating an attention mechanism in the execution unit, the out-of-order execution of multiple warps causes the operations of multiple warps to be intertwined, resulting in a low operation efficiency of the attention mechanism. Summary of the Invention
[0005] The present invention provides a method for operating an attention mechanism, an electronic device, and a storage medium to solve the defect that the out-of-order execution of multiple warps in related technologies results in a low operation efficiency of the attention mechanism.
[0006] The present invention provides a method for operating an attention mechanism, including: At a tensor calculation core, controlling two warps to alternately enter a tensor operation state and a waiting state, where the tensor operation state is used to continuously run a second tensor operation of one attention operation and a first tensor operation of another attention operation; And, after the first tensor operation ends, controlling a vector calculation core to run a vector operation, where the attention operation includes a first tensor operation, a vector operation, and a second tensor operation that are sequentially executed.
[0007] According to the method for operating an attention mechanism provided by the present invention, the controlling two warps to alternately enter a tensor operation state and a waiting state includes: For each of the two warps, after the warp enters the tensor operation state and before running the first tensor operation, setting the instruction priority of the warp to a first priority, and after the warp completes the second tensor operation, controlling the warp to issue a running instruction for the first tensor operation under the first priority to trigger the running of the first tensor operation.
[0008] According to an operation method of an attention mechanism provided by the present invention, controlling two warps to alternately enter a tensor operation state and a waiting state further includes: For each of the two warps, after the warp completes the first tensor operation, set the instruction priority of the warp to a second priority, where the second priority is lower than the first priority.
[0009] According to an operation method of an attention mechanism provided by the present invention, controlling two warps to alternately enter a tensor operation state and a waiting state includes: Based on the synchronization mechanism between the two warps, after one warp completes the first tensor operation, control the other warp to perform the second tensor operation.
[0010] According to an operation method of an attention mechanism provided by the present invention, after the first tensor operation ends, controlling a vector calculation core to perform a vector operation includes: At the vector calculation core, control each of the two warps to perform a vector operation after the first tensor operation of the warp ends.
[0011] According to an operation method of an attention mechanism provided by the present invention, before the vector operation of the warp is completed, at the tensor calculation core, control the warp to continuously be in the waiting state.
[0012] According to an operation method of an attention mechanism provided by the present invention, the sum of the time taken for the first tensor operation and the second tensor operation is greater than or equal to the time taken for the vector operation.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the operation method of the attention mechanism as described in any one of the above is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the operation method of the attention mechanism as described in any one of the above is implemented.
[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the operation method of the attention mechanism as described in any one of the above is implemented.
[0016] The method for operating an attention mechanism, electronic device, and storage medium provided by the present invention control two warps to alternately enter a tensor operation state and a waiting state at the tensor calculation core, and continuously perform the second tensor operation of an attention operation and the first tensor operation of another attention operation in the tensor operation state, avoiding the problem of interweaving of tensor calculations of the two warps and effectively improving the tensor calculation efficiency; and, after the first tensor operation ends, control the vector calculation core to perform a vector operation. The time consumed by the vector operation coincides with the time when the warp is in the waiting state in the tensor calculation core, and the vector operations of the two warps at the vector calculation core also alternate, thereby improving the efficiency of vector calculation and further improving the overall efficiency of the operation of the attention mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions in the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a schematic diagram of the interaction between a tensor calculation core and a vector calculation core in the related art; Figure 2 is a schematic flowchart of the flash attention algorithm in the related art; Figure 3 is a schematic diagram of the operation of the attention mechanism in the related art; Figure 4 is one of the schematic flowcharts of the method for operating the attention mechanism provided by the present invention; Figure 5 is another schematic flowchart of the method for operating the attention mechanism provided by the present invention; Figure 6 is a schematic structural diagram of the device for operating the attention mechanism provided by the present invention; Figure 7 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0020] With the development of artificial intelligence technology, artificial intelligence chips are often used in general computing or artificial intelligence computing. The artificial intelligence chips here can be Graphics Processing Units (GPUs), or they can also be NPUs (Neural network Processing Units), DPUs (Deep learning Processing Units), APUs (Accelerated Processing Units), and GPGPUs (General-Purpose computing on Graphics Processing Unit), etc.
[0021] As the basic unit for executing computing tasks in an artificial intelligence chip, the execution unit can activate multiple warps simultaneously through the hardware multithreading mechanism, thereby supporting the parallel execution of multiple warps. Here, the warp, which is the scheduling unit of the execution unit, is a group composed of threads that execute in parallel. For example, each warp can contain 32 threads. During the calculation process, the execution unit can decompose the computing tasks into warps and allocate the warps to the available computing cores within the execution unit to support the parallelism of the warps through the computing resources at the computing cores. The parallelism of multiple warps can mutually cover the idle time in instruction execution, thereby improving the overall execution efficiency of the execution unit.
[0022] Furthermore, the computing cores within the execution unit can include Tensor Cores (Tcores) and Vector Cores (Vcores). Figure 1 It is a schematic diagram of the interaction between the Tensor Core and the Vector Core in the related technology. As Figure 1 shown, in an artificial intelligence chip, the input data is transmitted from the memory or cache to the Tensor Core along at least one data transmission channel, and the tensor calculation is completed in the Tensor Core to obtain the tensor calculation result. Then, the tensor calculation result is transmitted to the memory or cache, or the Vector Core, as the output data.
[0023] However, the emission of the operation running instructions of the warp is restricted by the computing resources, such as the number of Arithmetic Logic Units (ALUs). It can be understood that the operation of an operation requires matching corresponding computing resources. Therefore, the operation running instructions of a warp need to occupy a specific type of computing unit. If the operation running instructions of different warps need to occupy the same computing unit, then only the operation running instructions of one warp can be selected for emission at the same moment.
[0024] Therefore, in the parallel scenario of multiple warps, at the same moment, the execution unit can only control a few warps to issue arithmetic operation instructions (usually 1 instruction or the arithmetic operation instructions of 2 warps). In the current artificial intelligence chip design, at a certain moment, usually the warps that meet the instruction issue conditions execute instruction issue. If there are multiple warps that meet the instruction issue conditions at the same moment, then 1 or 2 warps need to be selected from multiple warps to execute instruction issue according to the pre-set rules. The rules mentioned here can be round-robin issue.
[0025] It can be imagined that in the case of parallel multiple warps in the execution unit, due to the limitation of the number of instructions issued at the same moment, the order in which multiple warps issue arithmetic operation instructions does not necessarily match the execution order of each arithmetic operation allocated to multiple warps in advance for realizing the computing task. This results in that multiple warps may be out-of-order executed in the execution unit.
[0026] However, in some computing scenarios, multiple warps executing in a fixed order may achieve higher computing efficiency compared to out-of-order execution.
[0027] For example, the Attention mechanism, as an important part of the Large Language Model (LLM), is currently usually implemented based on the flash attention algorithm. Figure 2 It is a schematic flow diagram of the flashattention algorithm in the related technology. As Figure 2 shown, in the flash attention algorithm, the input data for the attention mechanism operation can be obtained first, that is, three vectors: Q (Query), K (Key), and V (Value). For Q and K T (the transpose of K), perform the QK T operation. For the result of the QK T operation, perform the Softmax operation, and then perform the QKV operation on the result of the Softmax operation and V, thereby obtaining the final output O.
[0028] It can be seen that the main body of the flash attention algorithm consists of three parts of operations, namely QK T , Softmax, and QKV. Among them, QK T , QKV are matrix operations, usually implemented by the Tensor Core (Tcore) in the execution unit, and Softmax is a vector operation, usually implemented by the Vector Core (Vcore) in the execution unit.
[0029] As can be seen from Figure 2 , during the operation of the attention mechanism, QK T , Softmax, and QKV are executed in sequence. Correspondingly, within a warp, QK T , Softmax, and QKV are also executed in sequence. That is, for a warp, during the operation of the tensor computing core, the vector computing core is in a waiting state, and vice versa, during the operation of the vector computing core, the tensor computing core is in a waiting state.
[0030] To improve the execution efficiency, two warps can be executed in parallel. In the case of two warps being executed in parallel, ideally, when the two warps are executed in a fixed order, it is that while one warp is running on the tensor computing core, the other warp is running on the vector computing core, and there is no situation where the two warps are running on the tensor computing core simultaneously. However, limited by the number of instructions issued at the same time, the two warps may actually be executed out of order. Specifically, when one warp is running on the tensor computing core, the other warp may also transfer to the tensor computing core to run. That is, there is a situation where the two warps are running on the tensor computing core simultaneously. In other words, the operations of the two warps on the tensor computing core and the vector computing core are intertwined.
[0031] Figure 3 is a schematic diagram of the operation of the attention mechanism in the related technology. As Figure 3 shown, warps 1 and 2 execute the attention operation in parallel, where the white boxes represent the operations executed at the same tensor computing core, and the boxes filled with slashes represent the operations executed at the same vector computing core. As can be seen from Figure 3 , when warp 1 executes QK T (i.e., QK Figure 3 in T (i + 2)) of the (i + 2)-th attention operation, it is intertwined with QKV (i.e., QKV(j + 1) in Figure 3 ) of the (j + 1)-th attention operation executed by warp 2 on the tensor computing core. That is, there is a parallel situation of QK T (i + 2) of warp 1 and QKV(j + 1) of warp 2 on the tensor computing core. In addition, when warp 1 executes QKV (i.e., QKV(i + 2) in Figure 3 ) of the (i + 2)-th attention operation, it is intertwined with QK T (i.e., QK Figure 3 in T (j + 2)) of the (j + 2)-th attention operation executed by warp 2 on the tensor computing core. That is, there is a situation where QKV(i + 2) of warp 1 and QKT (j + 2) parallel case. This intertwined situation on the same tensor computing core will lead to a shortage of computing resources for a period of time. The computing efficiency of both intertwined operations will decrease, the computing time will be lengthened, resulting in a low operating efficiency of the attention mechanism.
[0032] And, as Figure 3 can be seen, even if warp 1 and warp 2 are executed in parallel, there are still some idle times on the tensor computing core and the vector computing core. For example, the prerequisite for warp 1 to execute the Softmax of the (i + 1)-th attention operation is that the QK of the (i + 1)-th attention operation T (i + 1) is completed. However, since the QK of warp 1 T (i + 1) is intertwined with the QKV(j) of warp 2 at the same tensor computing core, this causes the computing time of the QK of warp 1 T (i + 1) to be lengthened, and the start time of the Softmax of the (i + 1)-th attention operation is also delayed accordingly. The vector computing core used to execute the Softmax is idle for a long time after completing the Softmax of the j-th attention operation of warp 2, which also causes a certain degree of waste of computing resources and further reduces the operating efficiency of the attention mechanism.
[0033] In view of the above situation, an embodiment of the present invention provides a method for operating an attention mechanism. Figure 4 is one of the flow diagrams of the method for operating an attention mechanism provided by the present invention. As Figure 4 shown, this method is applied to any artificial intelligence chip, and specifically can be applied to the execution unit of an artificial intelligence chip. The method includes: Step 410, at the tensor computing core, control two warps to alternately enter the tensor operation state and the waiting state. The tensor operation state is used to continuously run the second tensor operation of one attention operation and the first tensor operation of another attention operation.
[0034] Step 420, after the first tensor operation ends, control the vector computing core to run the vector operation. The attention operation includes a first tensor operation, a vector operation, and a second tensor operation that are executed sequentially.
[0035] Specifically, in the execution unit of the AI chip, a tensor calculation core Tcore and a vector calculation core Vcore can be set up. Among them, the tensor calculation core is a hardware unit dedicated to accelerating tensor operations, and the vector calculation core is a hardware unit dedicated to accelerating vector operations. It should be noted that in the embodiments of the present invention, for an execution unit in the AI chip, there is only one tensor calculation core and one vector calculation core. The following control operations performed on the tensor calculation core are all for the same tensor calculation core, and the control operations performed on the vector calculation core are all for the same vector calculation core.
[0036] During the operation of the attention mechanism, multiple attention operations can be performed. Each attention operation can include three operations executed in sequence, namely the first tensor operation, the vector operation, and the second tensor operation. Among them, the first tensor operation is the QK T operation, the vector operation is the Softmax operation, and the second tensor operation is the QKV operation. Moreover, the first tensor operation and the second tensor operation can be allocated to the tensor calculation core for operation, while the vector operation can be allocated to the vector calculation core for operation. For any warp in the execution unit, the process of completing an attention operation through this warp can be that the warp first runs the first tensor operation in the tensor calculation core, then runs the vector operation in the vector calculation core, and finally runs the second tensor operation in the tensor calculation core.
[0037] In the embodiments of the present invention, to improve the operation efficiency of the attention mechanism, two warps are executed in parallel in the execution unit. And to avoid the out-of-order problem when two warps are in parallel in the related art, in the embodiments of the present invention, for the operation of the attention mechanism at the tensor calculation core, the two warps are controlled to alternately enter the tensor operation state and the waiting state.
[0038] Here, controlling the two warps to alternately enter the tensor operation state and the waiting state during the operation at the tensor calculation core specifically means that at the tensor calculation core, when one warp enters the tensor operation state, the other warp is controlled to enter the waiting state, and when one warp is in the waiting state, the other warp is controlled to enter the tensor operation state. Thus, at the tensor calculation core, there is no situation where two warps are simultaneously in the tensor operation state, that is, there is no problem of tensor operation interleaving of two warps at the tensor calculation core.
[0039] Further, for each warp, in the tensor operation state, the warp performs tensor operations in the attention operation at the tensor computing core, specifically including the first tensor operation and the second tensor operation. Moreover, considering that in one attention operation, it is executed in the order of the first tensor operation, vector operation, and second tensor operation, that is, after the first tensor operation ends, it is necessary to wait for the vector operation to complete before the second tensor operation can be executed. In multiple attention operations, the second tensor operation of the previous attention operation and the first tensor operation of the next attention operation are independent of each other, and the second tensor operation of the previous attention operation and the first tensor operation of the next attention operation can run continuously without waiting for other operation results additionally.
[0040] Therefore, in the embodiment of the present invention, the tensor operation state is defined as continuously running the second tensor operation of one attention operation and the first tensor operation of another attention operation. Assuming that one of the attention operations is the i-th attention operation and the other attention operation is the (i + 1)-th attention operation, the tensor operation state can be the continuous running of QKV(i) and QK T (i + 1). Where QKV(i) is the second tensor operation of the i-th attention operation, and QK T (i + 1) is the first tensor operation of the (i + 1)-th attention operation.
[0041] For each warp, the waiting state refers to the state where the warp waits at the tensor computing core and does not perform tensor operations for the time being. It can be understood that in the waiting state, the warp is in an idle state at the tensor computing core.
[0042] In addition, for each warp, vector operations can be run at the vector computing core. Moreover, the vector operation is run after the first tensor operation ends, and the first tensor operation is the latter of the two tensor operations run by the warp in the tensor operation state. That is, for a warp, when the warp ends the tensor operation state and enters the waiting state at the tensor computing core, it runs the vector operation at the vector computing core.
[0043] For example, Figure 5 is the second flow chart showing the operation method of the attention mechanism provided by the present invention. As Figure 5 shown, the two warps can be warp 1 and warp 2. The operations of the two warps at the tensor computing core are represented by white boxes, and the operations at the vector computing core are represented by boxes filled with slashes. By Figure 5It can be seen that after warp 1 enters the tensor operation state, warp 2 is in a waiting state in the tensor calculation core and runs a vector operation (Softmax) in the vector calculation core. Moreover, in the tensor operation state of warp 1, after the QKV(i) operation is completed, warp 2 does not insert a tensor operation instruction, but warp 1 continues to perform the QK T (i + 1) operation, thus realizing the continuous operation of QKV(i) and QK T (i + 1). When warp 1 completes the tensor operation state in the tensor calculation core and enters the waiting state, warp 1 runs a vector operation in the vector calculation core, and at the same time, warp 2 transfers from the waiting state to the tensor operation state. Moreover, in the tensor operation state of warp 2, after the QKV(j) operation is completed, warp 1 does not insert a tensor operation instruction, but warp 2 continues to perform the QK T (j + 1) operation, thus realizing the continuous operation of QKV(j) and QK T (j + 1). When warp 2 completes the tensor operation state in the tensor calculation core and enters the waiting state, warp 2 runs a vector operation in the vector calculation core, and at the same time, warp 1 transfers from the waiting state to the tensor operation state, and so on. The two warps alternately enter the tensor operation state and the waiting state at the tensor calculation core, and each warp runs a vector operation in the vector calculation core while entering the waiting state at the tensor calculation core.
[0044] It can be understood that in the above method for running the attention mechanism, the two warps alternately enter the tensor operation state and the waiting state at the tensor calculation core, which can effectively avoid the situation where the two warps are intertwined in the tensor calculation core. Moreover, when one warp enters the waiting state in the tensor calculation core, the other warp enters the tensor operation state in the tensor calculation core. Overall, the running time of the tensor calculation core can be filled by the two warps and there is basically no idle time, thus greatly improving the resource utilization rate of the tensor calculation core and helping to improve the calculation efficiency of the attention mechanism in tensor operations.
[0045] In addition, each warp runs a vector operation in the vector calculation core while entering the waiting state at the tensor calculation core. The time of the vector operation on the warp basically coincides with the time when the warp is in the waiting state in the tensor calculation core. Thus, for a single attention operation, the efficiency of the attention operation is guaranteed. Moreover, the vector operations on the two warps are also alternately performed at the vector calculation core. Overall, the running time of the vector calculation core is also basically filled by the two warps, and the resource utilization rate of the vector calculation core is also improved.
[0046] It should be noted that the attention mechanism in the embodiments of the present invention can be a component of the model. The method for operating the attention mechanism provided by the embodiments of the present invention can be applied to at least one of the model training process and the model application process. The model referred to here can be applied to different processing scenarios such as speech, image, text, video, etc. For example, in the speech processing scenario, the model containing the attention mechanism can be a speech recognition model, a voiceprint recognition model, etc.; in the image processing scenario, the model containing the attention mechanism can be a face recognition model, an object recognition model, an object segmentation model, etc.; in the text processing scenario, the model containing the attention mechanism can be a text translation model, a text correction model, a dialogue model, etc.; in the video processing scenario, the model containing the attention mechanism can be a video generation model, an object tracking model, etc. The embodiments of the present invention do not make specific limitations on this.
[0047] In the method provided by the embodiments of the present invention, by controlling two warps to alternately enter the tensor operation state and the waiting state at the tensor computing core, and continuously running the second tensor operation of one attention operation and the first tensor operation of another attention operation in the tensor operation state, the problem of interweaving of tensor calculations of the two warps is avoided, effectively improving the tensor calculation efficiency; and, after the first tensor operation ends, controlling the vector computing core to run vector operations, the time consumed by the vector operations coincides with the time when the warps are in the waiting state at the tensor computing core, and the vector operations of the two warps at the vector computing core also alternate, so that the efficiency of vector calculation is also improved, and thus the overall efficiency of the operation of the attention mechanism is improved.
[0048] It can be understood that, by improving the overall efficiency of the operation of the attention mechanism, the method provided by the embodiments of the present invention can effectively improve the efficiency of the model containing the attention mechanism during training and inference, and thus ensure the operation efficiency of the model in various processing scenarios.
[0049] Based on the above embodiments, in step 410, the controlling the two warps to alternately enter the tensor operation state and the waiting state includes: For each of the two warps, after the warp enters the tensor operation state and before running the first tensor operation, set the instruction priority of the warp to the first priority, and after the warp completes the second tensor operation, control the warp to issue the running instruction of the first tensor operation under the first priority to trigger the running of the first tensor operation.
[0050] Specifically, in order to enable two warps to alternately enter the tensor operation state and the waiting state, and to avoid the situation where after a warp completes the second tensor operation in the tensor operation state, another warp's tensor operation is inserted out of order, in the embodiments of the present invention, a priority is set for the execution of the first tensor operation, thereby ensuring that for each warp, continuous execution of the second tensor operation and the first tensor operation can be achieved.
[0051] At the tensor calculation core, for each of the two warps, after the warp enters the tensor operation state and specifically before running the first tensor operation in the tensor operation state, that is, during the process of the warp running the second tensor operation or when the warp completes the second tensor operation, the instruction priority of the warp can be set to the first priority. It can be understood that the first priority here is higher than the instruction priority of the other warp.
[0052] Thus, after the warp completes the second tensor operation, the warp is in a condition to issue the running instruction of the first tensor operation. Assume that at the same time, another warp also meets the condition of entering the tensor operation state and issuing the running instruction of the second tensor operation. Then, in the execution unit, there is a need to issue two running instructions at the same time, but the execution unit itself only supports issuing one instruction at a time. At this time, the execution unit can compare the instruction priorities of the two warps and preferentially control the warp with the higher instruction priority to issue the instruction.
[0053] Since the warp in the tensor operation state has set its instruction priority to the first priority, which is higher than the instruction priority of the other warp, before running the first tensor operation, after the second tensor operation is completed, because the instruction priority of this warp is higher than that of the other warp, this warp can be controlled to issue the running instruction of the first tensor operation, while the other warp cannot issue the running instruction of the second tensor operation due to its lower instruction priority and remains in the waiting state without entering the tensor operation state.
[0054] In the embodiments of the present invention, by setting the instruction priority of the warp to the first priority after the warp enters the tensor operation state and before running the first tensor operation, it is possible to avoid the situation where a warp is inserted with a running instruction by another warp during the process of the warp being in the tensor operation state, resulting in the interweaving of the two warps in the tensor calculation core. Thus, it is ensured that each warp can complete and continuously execute the second tensor operation of an attention operation and the first tensor operation of another attention operation in the tensor operation state, thereby effectively improving the tensor calculation efficiency under the attention mechanism.
[0055] Based on any of the above embodiments, in step 410, the controlling the two warps to alternately enter the tensor operation state and the waiting state further includes: For each of the two warps, after the warp completes the first tensor operation, set the warp instruction priority to a second priority, where the second priority is lower than the first priority.
[0056] Specifically, to ensure that after one warp completes the tensor operation state, the other warp can quickly switch from the waiting state to the tensor operation state. At the tensor calculation core, for each of the two warps, after the warp completes the first tensor operation, that is, after the warp exits the tensor operation state, the instruction priority of the warp can be set to the second priority. It can be understood that the second priority here is lower than the first priority in the above embodiment. Thus, for this warp, it is equivalent to restoring the instruction priority that was increased when the warp was in the tensor operation state to the original instruction priority after the tensor operation state.
[0057] In this way, for each warp, at the tensor calculation core, during the process of the warp running the first tensor operation in the tensor operation state, the instruction priority of the warp always remains at the first priority, which is higher than the instruction priority of the other warp. Therefore, during the process of the warp performing the first tensor operation state in the tensor operation state, even if the other warp meets the condition for issuing the running instruction of the tensor operation, it will not be able to issue the running instruction because the instruction priority is lower than the first priority. Thus, the other warp remains in the waiting state.
[0058] And after one warp switches from the tensor operation state to the waiting state, adjust the instruction priority of the warp to the second priority. At this time, the instruction priority of this warp is the same as that of the other warp (both are the second priority). Since after the first tensor operation is completed, this warp needs to wait for the vector operation to complete before switching back to the tensor operation state, this warp does not meet the condition for issuing instructions during the process of waiting for the vector operation to complete. While the vector operation of the other warp on the vector calculation core has been completed, and the other warp already meets the condition for issuing the running instruction of the tensor operation. Therefore, it is possible to control the other warp to issue the running instruction of the second tensor operation to trigger the second tensor operation of the other warp in the tensor calculation core. Thus, the other warp can switch from the waiting state to the tensor operation state.
[0059] For example, in Figure 5 , during the process of warp 1 running QKV(i), the instruction priority of warp 1 at the tensor calculation core can be increased from the second priority to the first priority. Thus, after the QKV(i) operation is completed, since the instruction priority of warp 1 is the first priority and the instruction priority of warp 2 is the second priority, and the instruction priority of warp 1 is higher, warp 1 can issue QKT The execution instruction of (i + 1). At this time, the execution instruction of warp 2 cannot be issued, and QKV(i) and QK can be implemented on warp 1 T The continuous execution of (i + 1). And, in QK T After the execution of (i + 1) is completed, the instruction priority of warp 1 in the tensor computing core can be restored from the first priority to the second priority. At this time, warp 2 can issue the execution instruction of QKV(j).
[0060] Through the above method of setting instruction priority, the resource utilization rate of the attention mechanism running in the tensor computing core can be increased from 70% to more than 90%, which has a huge improvement effect on the running efficiency of large language models.
[0061] In the embodiment of the present invention, by setting the instruction priority of the warp to the second priority after the warp completes the first tensor operation, another warp can quickly switch from the waiting state to the tensor operation state, thereby realizing the alternation of the two warps between the tensor operation state and the waiting state, and thus effectively improving the tensor computing efficiency under the attention mechanism.
[0062] Based on any of the above embodiments, in step 410, controlling the two warps to alternately enter the tensor operation state and the waiting state includes: Based on the synchronization mechanism between the two warps, after one warp completes the first tensor operation, controlling the other warp to execute the second tensor operation.
[0063] Specifically, in order to realize the alternation of the two warps entering the tensor operation state and the waiting state and avoid the situation that after one warp completes the second tensor operation in the tensor operation state, it is inserted out of order into the tensor operation of the other warp. In the embodiment of the present invention, a synchronization mechanism is set between the two warps.
[0064] It can be understood that the role of the synchronization mechanism between the two warps is to coordinate the execution order of the two warps, thereby ensuring the sequential execution of the two warps and facilitating the out-of-order execution of the parallel warps. In the embodiment of the present invention, the completion of the first tensor operation by one warp can be set as the execution condition for the other warp to execute the second tensor operation. Thus, at the tensor computing core, during the process of one warp running the first tensor operation in the tensor operation state, the synchronization mechanism controls the other warp to wait for the completion of the first tensor operation of this warp, and after the completion of the first tensor operation of this warp, controls the other warp to enter the tensor operation state to start running the second tensor operation.
[0065] Thus, through the synchronization mechanism between two warps, the replacement of the two warps in the tensor operation state and the waiting state is realized, thereby effectively improving the tensor calculation efficiency under the attention mechanism.
[0066] Based on any of the above embodiments, after the first tensor operation is completed, controlling the vector calculation core to perform vector operations includes: At the vector calculation core, controlling each warp in the two warps to perform vector operations after the first tensor operation of the warp is completed.
[0067] Specifically, two warps parallel in the execution unit can both perform vector operations at the vector calculation core. And for each warp in the two warps, after the first tensor operation is completed in the tensor calculation core, that is, after switching from the tensor calculation state to the waiting state in the tensor calculation core, the warp can perform vector operations at the vector calculation core. The result of the vector operation obtained thereby is the data used for the second tensor operation when the warp switches back from the waiting state to the tensor calculation state next time.
[0068] Since for each warp, the time for the warp to perform vector operations at the vector calculation core basically coincides with the time in the waiting state in the tensor calculation core, and the two warps alternately enter the waiting state at the tensor calculation core, therefore at the vector calculation core, the time for the two warps to perform vector operations basically does not coincide, which helps to improve the calculation efficiency of the attention mechanism in vector operations.
[0069] Based on any of the above embodiments, the method further includes: Before the vector operation of the warp is completed, at the tensor calculation core, controlling the warp to continuously be in the waiting state.
[0070] Specifically, in one attention operation, the first tensor operation, the vector operation, and the second tensor operation are executed in sequence. Thus, for one warp, when the warp switches from the tensor operation state to the waiting state, the first tensor operation has been completed. Thus, the warp performs vector operations at the vector calculation core and uses the result of the vector operation as the input for the second tensor operation.
[0071] Based on this, before the vector operation of the warp is completed, that is, before the result of the vector operation is obtained, the warp cannot obtain the input for the second tensor operation and thus cannot enter the tensor operation state to perform the second tensor operation. At this time, the warp can be controlled to continuously be in the waiting state until another warp switches from the tensor operation state to the waiting state and the vector operation of this warp is completed.
[0072] Based on any of the above embodiments, the sum of the time consumption of the first tensor operation and the second tensor operation is greater than or equal to the time consumption of the vector operation.
[0073] Specifically, in most cases, from the perspective of the computing power of the tensor calculation core and the vector calculation core in the execution unit, in one attention operation, the time consumption of the vector operation is usually less than or equal to the sum of the time consumption of the first tensor operation and the second tensor operation, that is, the time consumption of the vector operation is less than or equal to the time consumption when the warp is in the tensor operation state.
[0074] Since two warps alternately enter the tensor operation state and the waiting state in the tensor calculation core, that is, when one warp is in the tensor operation state, the other warp is in the waiting state, the time consumption of the tensor operation state is the same as the time consumption of the waiting state.
[0075] Thus, the time consumption of a warp in the waiting state in the tensor calculation core can cover the time consumption of the warp running the vector operation in the vector calculation core, and the warp does not need to spend additional time waiting for the end of the vector operation. Generally speaking, in the tensor calculation core, when one warp cuts from the tensor operation state to the waiting state, the other warp cuts from the waiting state to the tensor operation state, and the running time of the tensor calculation core can be filled by the two warps and there is basically no idle time, which greatly improves the resource utilization rate of the tensor calculation core and helps to improve the calculation efficiency of the attention mechanism in tensor operations.
[0076] Next, the attention mechanism operation device provided by the present invention will be described. The attention mechanism operation device described below can be correspondingly referred to the attention mechanism operation method described above.
[0077] Figure 6 is a schematic structural diagram of the attention mechanism operation device provided by the present invention, as Figure 6 shown, the device includes: A tensor calculation unit 610, configured to control two warps to alternately enter the tensor operation state and the waiting state at the tensor calculation core, and the tensor operation state is used to continuously run the second tensor operation of one attention operation and the first tensor operation of another attention operation; A vector calculation unit 620, configured to control the vector calculation core to run a vector operation after the first tensor operation ends, and the attention operation includes a first tensor operation, a vector operation, and a second tensor operation executed in sequence.
[0078] The device provided by the embodiment of the present invention controls two warps to alternately enter the tensor operation state and the waiting state at the tensor calculation core, and continuously runs the second tensor operation of one attention operation and the first tensor operation of another attention operation in the tensor operation state, avoiding the problem of interlacing in the tensor calculations of the two warps and effectively improving the tensor calculation efficiency; and, after the first tensor operation is completed, it controls the vector calculation core to run vector operations. The time consumed by the vector operations coincides with the time when the warp is in the waiting state in the tensor calculation core, and the vector operations of the two warps at the vector calculation core also alternate, so that the efficiency of vector calculation is also improved, and thus the overall efficiency of the attention mechanism operation is improved.
[0079] Based on any of the above embodiments, the tensor calculation unit is specifically configured to: For each of the two warps, after the warp enters the tensor operation state and before running the first tensor operation, set the instruction priority of the warp to the first priority, and after the warp completes the second tensor operation, control the warp to issue the running instruction of the first tensor operation under the first priority to trigger the running of the first tensor operation.
[0080] Based on any of the above embodiments, the tensor calculation unit is further configured to: For each of the two warps, after the warp completes the first tensor operation, set the instruction priority of the warp to the second priority, and the second priority is lower than the first priority.
[0081] Based on any of the above embodiments, the tensor calculation unit is specifically configured to: Based on the synchronization mechanism between the two warps, after one warp completes the first tensor operation, control the other warp to run the second tensor operation.
[0082] Based on any of the above embodiments, the vector calculation unit is specifically configured to: At the vector calculation core, control each of the two warps to run vector operations after the first tensor operation of the warp is completed.
[0083] Based on any of the above embodiments, the tensor calculation unit is further configured to: Before the vector operation of the warp is completed, at the tensor calculation core, control the warp to continuously be in the waiting state.
[0084] Based on any of the above embodiments, the sum of the time consumed by the first tensor operation and the second tensor operation is greater than or equal to the time consumed by the vector operation.
[0085] Figure 7 Illustrates a schematic diagram of the physical structure of an electronic device, as Figure 7 shown. The electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logical instructions in the memory 730 to execute the method for running the attention mechanism. The method includes: At the tensor calculation core, control two warps to alternately enter the tensor operation state and the waiting state. The tensor operation state is used to continuously run the second tensor operation of one attention operation and the first tensor operation of another attention operation; And, after the first tensor operation ends, control the vector calculation core to run vector operations. The attention operation includes a first tensor operation, a vector operation, and a second tensor operation performed in sequence.
[0086] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, etc., which can store program codes.
[0087] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for running the attention mechanism provided by the above-mentioned various methods. The method includes: At the tensor calculation core, control two warps to alternately enter the tensor operation state and the waiting state. The tensor operation state is used to continuously run the second tensor operation of one attention operation and the first tensor operation of another attention operation; And, after the first tensor operation ends, control the vector calculation core to perform a vector operation. The attention operation includes a first tensor operation, a vector operation, and a second tensor operation that are sequentially executed.
[0088] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the attention mechanism operation method provided by the above various methods. The method includes: At the tensor calculation core, control two warps to alternately enter the tensor operation state and the waiting state. The tensor operation state is used to continuously execute the second tensor operation of one attention operation and the first tensor operation of another attention operation. And, after the first tensor operation ends, control the vector calculation core to perform a vector operation. The attention operation includes a first tensor operation, a vector operation, and a second tensor operation that are sequentially executed.
[0089] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0090] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0091] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for operating an attention mechanism, characterized in that: include: At the tensor computing core, controlling two thread warps to alternately enter a tensor operation state and a waiting state, wherein the tensor operation state is used to continuously run a second tensor operation of one attention operation and a first tensor operation of another attention operation; And, after the first tensor operation is completed, the vector calculation core is controlled to run the vector operation, and the attention operation includes the first tensor operation, the vector operation and the second tensor operation executed sequentially.
2. The attention mechanism operation method according to claim 1, characterized in that: The controlling the two thread warps to alternately enter the tensor operation state and the waiting state includes: For each of the two thread warps, after the thread warp enters the tensor operation state and before executing the first tensor operation, the instruction priority of the thread warp is set to the first priority level, and after the thread warp completes the second tensor operation, the thread warp is controlled to emit the execution instruction of the first tensor operation at the first priority level to trigger the execution of the first tensor operation.
3. The attention mechanism operation method according to claim 2, characterized in that: The controlling the two thread warps to alternately enter the tensor operation state and the waiting state also includes: For each of the two thread warps, after the thread warp completes the first tensor operation, setting the thread warp instruction priority to a second priority, where the second priority is lower than the first priority.
4. The attention mechanism operation method according to claim 1, characterized in that: The controlling the two thread warps to alternately enter the tensor operation state and the waiting state includes: Based on the synchronization mechanism between the two warps, after one warp completes the first tensor operation, another warp is controlled to execute the second tensor operation.
5. The attention mechanism operation method according to any one of claims 1 to 4, characterized in that: After the first tensor operation is completed, controlling the vector computing core to perform vector operations includes: At the vector computing core, each of the two warps is controlled to perform a vector operation after a first tensor operation of the warp is completed.
6. The attention mechanism operation method according to claim 5, characterized in that: Also includes: Before the vector operation of the thread warp is completed, at the tensor computing core, the thread warp is controlled to remain in the waiting state.
7. The attention mechanism operation method according to any one of claims 1 to 4, characterized in that: The sum of the time consumed by the first tensor operation and the second tensor operation is greater than or equal to the time consumed by the vector operation.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, it implements the attention mechanism operation method as described in any one of claims 1 to 7.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the attention mechanism operation method as described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the attention mechanism operation method as described in any one of claims 1 to 7.
Citation Information
Cited By
Attention mechanism calculation method and device, medium and product
CN121052309A
An attention mechanism calculation method, device, medium and product
CN121052309B