Implementation device and method for sharing register block by general-purpose computing core and tensor core
By introducing tensor core read request queues and life cycle counters in GPGPU design, the scheduler provides higher priority for general computing cores, solving the problem of insufficient register group read bandwidth and improving overall performance.
Patent Information
- Application Number
- CN202510461109.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
AI Technical Summary
In GPGPU design, when the general computing core shares the register group with the tensor core, the read bandwidth of the register group is insufficient, resulting in a degradation of overall performance. Especially when the tensor core reads operands, the general computing core cannot read at the same time, resulting in idle execution units.
By introducing tensor kernel read request queues, life cycle counters and clock cycle counters, the scheduler gives the general computing core higher priority without blocking tensor kernel read requests, maximizes the use of the read bandwidth of the register group, and uses tensor kernel operands to cache and schedule read requests.
Without affecting the read request of the tensor kernel, the reading efficiency of the general computing core is improved, the idle time is reduced, and the read bandwidth of the register group is maximized, so as to realize the full load operation between the general computing core and the tensor kernel.
Smart Images

Figure CN120295670A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence chips, and particularly to an implementation device and method for sharing a register file between a general computing core and a tensor core. Background Art
[0002] A GPGPU (General-purpose computing on graphics processing units) is a graphics processing unit that processes general computing tasks originally processed by a central processing unit. A register file is the storage hierarchy closest to the GPGPU. To support multiple warps simultaneously and meet the zero-overhead switching of warps, the capacity of the register file of the GPGPU is usually relatively large. In current popular GPGPU designs, a general computing core (vector core) and a tensor core share the register file, that is, when the general computing core and the tensor core process computing tasks, they both read operands from the register file and write the computing results back to the register file.
[0003] However, the read bandwidth of the register file is a limited resource. When the general computing core and the tensor core need to process computing tasks simultaneously, the register file does not always provide sufficient read bandwidth. Especially in some implementations, when the tensor core reads operands from the register file, the general computing core cannot simultaneously read operands from the register file, which will cause the execution units of the general computing core to be idle, resulting in an obvious overall performance degradation. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide an implementation device and method for sharing a register file between a general computing core and a tensor core, which can give a higher priority to the general computing core without blocking the read requests of the tensor core when the general computing core and the tensor core share the register file, maximize the utilization of the read bandwidth of the register file, and thus effectively improve the overall performance.
[0005] To achieve the above purpose, the embodiments of the present invention provide an implementation device for sharing a register file between a general computing core and a tensor core, which is applied to the tensor core and includes:
[0006] A tensor core read request unit, configured to issue a read request and store it in the tensor core read request queue, and at the same time, trigger the life cycle counter and the clock cycle counter to start counting;
[0007] A life cycle counter, configured to record the life cycle value of each read request in the tensor core read request queue, and the life cycle value decreases as the life cycle counter counts;
[0008] A tensor core read request scheduler, which is used to receive read requests and their lifecycle values from the tensor core read request queue, and when the lifecycle value of the to-be-executed read request reaches a preset value, emit the to-be-executed read request to a register bank to read tensor core operands from the register bank and store them in a tensor core operand cache; wherein, the to-be-executed read request is the read request at the exit of the tensor core read request queue; the priority of the general computing core occupying the register bank is higher than that of the tensor core before the lifecycle value of the to-be-executed read request reaches the preset value;
[0009] A clock cycle counter, which is used to send a stack-out signal to the tensor core operand cache when the count value reaches a preset trigger value, so as to trigger the tensor core operand cache to send tensor core operands to a tensor core execution unit.
[0010] Further, the lifecycle value is set with an initial value, the preset trigger value is the count value corresponding to when the clock cycle counter accumulatively counts up to N clock cycles, and the accumulative count value of the lifecycle value decreasing from the initial value to the preset value does not exceed N.
[0011] Further, the tensor core read request scheduler is further used for:
[0012] Sending the lifecycle value of the to-be-executed read request to a general computing core scheduling unit; wherein, the general computing core scheduling unit is used to process read requests of the general computing core according to the lifecycle value of the to-be-executed read request.
[0013] Further, the general computing core scheduling unit processes read requests of the general computing core according to the lifecycle value of the to-be-executed read request, which specifically includes:
[0014] Judging whether a read request of the general computing core can be emitted to the register bank according to the lifecycle value of the to-be-executed read request;
[0015] When the lifecycle value of the to-be-executed read request is not the preset value, it is determined that a read request of the general computing core can be emitted to the register bank;
[0016] When the lifecycle value of the to-be-executed read request is the preset value, it is determined that a read request of the general computing core cannot be emitted to the register bank.
[0017] Further, the general computing core scheduling unit is further used for:
[0018] When it is determined that a read request of the general computing core can be emitted to the register bank, determining the maximum number of general computing core operands that can be currently read in the read request of the general computing core according to the current value of the lifecycle value of the to-be-executed read request.
[0019] Furthermore, the tensor core read request scheduler is further configured to:
[0020] Receive the read operand information sent by the general computing core scheduling unit;
[0021] Process the read requests in the tensor core read request queue according to the value of the read operand information; wherein, the value of the read operand information is used to indicate whether the register bank is occupied by the general computing core.
[0022] Furthermore, the tensor core read request scheduler processes the read requests in the tensor core read request queue according to the value of the read operand information, specifically including:
[0023] Judge whether the read requests in the tensor core read request queue can be issued to the register bank according to the value of the read operand information;
[0024] When the value of the read operand information is the first preset value, judge whether the life cycle value of the to-be-executed read request is the preset value; wherein, the first preset value is used to indicate that the register bank has been occupied by the general computing core;
[0025] If so, it is determined that the read requests in the tensor core read request queue can be issued to the register bank;
[0026] If not, it is determined that the read requests in the tensor core read request queue cannot be issued to the register bank.
[0027] Furthermore, the tensor core read request scheduler processes the read requests in the tensor core read request queue according to the value of the read operand information, and further includes:
[0028] When the value of the read operand information is the second preset value, it is determined that the read requests in the tensor core read request queue can be issued to the register bank; wherein, the second preset value is used to indicate that the register bank is not occupied by the general computing core.
[0029] Furthermore, the tensor core read request scheduler is further configured to:
[0030] When it is determined that the read requests in the tensor core read request queue can be issued to the register bank, issue the read requests in the tensor core read request queue to the register bank to read the tensor core operands from the register bank and store them in the tensor core operand cache.
[0031] To achieve the above object, an embodiment of the present invention further provides a method for implementing a general computing core and a tensor core sharing a register bank, which is applicable to the implementation device described in any one of the above, including:
[0032] A read request is sent through the tensor core read request unit and stored in the tensor core read request queue. Meanwhile, the life cycle counter and the clock cycle counter are triggered to start counting.
[0033] The life cycle value of each read request in the tensor core read request queue is recorded by the life cycle counter, and the life cycle value decreases as the life cycle counter counts.
[0034] The tensor core read request scheduler receives the read request and its life cycle value from the tensor core read request queue, and when the life cycle value of the to-be-executed read request reaches a preset value, the to-be-executed read request is emitted to the register bank to read the tensor core operand from the register bank and store it in the tensor core operand cache; wherein, the to-be-executed read request is the read request at the exit of the tensor core read request queue; the priority of the general computing core occupying the register bank is higher than that of the tensor core before the life cycle value of the to-be-executed read request reaches the preset value.
[0035] When the count value of the clock cycle counter reaches a preset trigger value, a stack-out signal is sent to the tensor core operand cache to trigger the tensor core operand cache to send the tensor core operand to the tensor core execution unit.
[0036] The embodiment of the present invention provides an implementation device and method for sharing a register bank between a general computing core and a tensor core. A read request is sent through the tensor core read request unit and stored in the tensor core read request queue. Meanwhile, the life cycle counter and the clock cycle counter are triggered to start counting. The life cycle value of each read request in the tensor core read request queue is recorded by the life cycle counter, and the life cycle value decreases as the life cycle counter counts. The tensor core read request scheduler receives the read request and its life cycle value from the tensor core read request queue, and when the life cycle value of the to-be-executed read request reaches a preset value, the to-be-executed read request is emitted to the register bank to read the tensor core operand from the register bank and store it in the tensor core operand cache, wherein the to-be-executed read request is the read request at the exit of the tensor core read request queue, and the priority of the general computing core occupying the register bank is higher than that of the tensor core before the life cycle value of the to-be-executed read request reaches the preset value. When the count value of the clock cycle counter reaches a preset trigger value, a stack-out signal is sent to the tensor core operand cache to trigger the tensor core operand cache to send the tensor core operand to the tensor core execution unit. The embodiment of the present invention can give a higher priority to the general computing core without blocking the read request of the tensor core when the general computing core and the tensor core share the register bank, maximize the utilization of the read bandwidth of the register bank, and thus effectively improve the overall performance. Description of the Drawings
[0037] Figure 1It is a schematic structural diagram of a preferred embodiment of an implementation device for a general computing core and a tensor core to share a register bank provided by the present invention;
[0038] Figure 2 It is a schematic flowchart of a preferred embodiment of a method for implementing a general computing core and a tensor core to share a register bank provided by the present invention. Detailed implementation manners
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art in the technical field without making creative efforts fall within the protection scope of the present invention.
[0040] It should be noted that in traditional GPGPU designs, the general computing core and the tensor core do not share the register bank when processing computing tasks. Generally, the general computing core monopolizes the register bank and reads operands from the register bank, while the tensor core reads operands from the tensor memory. This requires an additional tensor memory to be set up, and the general computing core and the tensor core usually need to exchange data. At this time, it is necessary to first load data from the tensor memory into the register bank, resulting in additional performance overhead.
[0041] Therefore, sharing the register bank between the general computing core and the tensor core has become the current mainstream trend. However, the read bandwidth of the register bank is a limited resource. When the general computing core and the tensor core need to process computing tasks simultaneously, they need to request to read operands from the register bank at the same time, which may exceed the upper limit of the read bandwidth provided by the register bank, resulting in a slowdown in data reading speed and thus affecting the execution progress of the computing tasks. Especially in some implementations, when the tensor core accesses the register bank, it will monopolize or preferentially occupy the read bandwidth of the register bank, making it impossible for the general computing core to perform read operations simultaneously. This is because the hardware design may set priorities or mutual exclusion controls for accessing the register bank in order to ensure that specific computing tasks of the tensor core (such as matrix multiplication in deep learning, etc.) can be completed efficiently. For example, when the tensor core performs large-scale matrix operations, it needs to quickly read a large amount of data. To avoid data transmission interruptions or errors, the hardware may prevent the general computing core from accessing the register bank at the same time, resulting in the idle execution units of the general computing core, wasting resources, and ultimately reducing the overall performance.
[0042] Exemplarily, assume that the tensor core requires a total of 10 clock cycles to process a certain computational task. It may only take 5 clock cycles to read the operands from the register file. However, since the read bandwidth of the register file is continuously and fully occupied by the tensor core for 5 clock cycles, during these 5 clock cycles, the general-purpose computing core cannot perform read operations simultaneously. Therefore, it will be idle for 5 clock cycles, which will cause the general-purpose computing core to be idle during these 5 idle clock cycles. The execution unit of the general-purpose computing core is idle, and the longer the idle time, the lower the overall performance.
[0043] To solve the above problems, an implementation device for sharing a register file between a general-purpose computing core and a tensor core is provided in an embodiment of the present invention. Refer to Figure 1 As shown, it is a schematic structural diagram of a preferred embodiment of an implementation device for sharing a register file between a general-purpose computing core and a tensor core provided by the present invention. The implementation device is applied to the tensor core and includes:
[0044] A tensor core read request unit, configured to issue a read request and store it in the tensor core read request queue. At the same time, it triggers the life cycle counter and the clock cycle counter to start counting;
[0045] A life cycle counter, configured to record the life cycle value of each read request in the tensor core read request queue, and the life cycle value decreases as the life cycle counter counts;
[0046] A tensor core read request scheduler, configured to receive the read request and its life cycle value from the tensor core read request queue, and when the life cycle value of the to-be-executed read request reaches a preset value, transmit the to-be-executed read request to the register file to read the tensor core operands from the register file and store them in the tensor core operand cache; wherein, the to-be-executed read request is the read request at the exit of the tensor core read request queue; the priority of the general-purpose computing core to occupy the register file is higher than that of the tensor core before the life cycle value of the to-be-executed read request reaches the preset value;
[0047] A clock cycle counter, configured to send a stack-out signal to the tensor core operand cache when the count value reaches a preset trigger value, to trigger the tensor core operand cache to send the tensor core operands to the tensor core execution unit.
[0048] It should be noted that the implementation device is applicable to the scenario where the general-purpose computing core and the tensor core share the register file. As Figure 1 shown, the implementation device can be set in the tensor core. The tensor core includes the implementation device and the tensor core execution unit. The implementation device is mainly composed of a tensor core read request unit, a tensor core read request queue, a life cycle counter, a clock cycle counter, a tensor core read request scheduler, and a tensor core operand cache.
[0049] Combined Figure 1 As shown, during the process of sharing the register bank between the general computing core and the tensor core, the working principle of the implementation device is as follows:
[0050] The tensor core read request unit is mainly used to issue read requests for the tensor core and temporarily store the issued read requests for the tensor core in the tensor core read request queue. At the same time, the tensor core read request unit is also used to, when issuing a read request for the tensor core, trigger the life cycle counter and the clock cycle counter to start independent counting respectively (for example, every time 1 clock cycle passes, the count value of the life cycle counter automatically increments by 1, and the count value of the clock cycle counter automatically increments by 1);
[0051] The life cycle counter is mainly used to record the life cycle value (i.e., life cnt) of each read request for the tensor core in the tensor core read request queue, and the life cycle value of each read request for the tensor core will decrease from the initial value as the life cycle counter counts (for example, every time 1 clock cycle passes, the count value of the life cycle counter automatically increments by 1, while life cnt automatically decrements by 1); among them, each read request for the tensor core has its own life cnt, which is used to mark its real-time life cycle, and life cnt in the tensor core read request queue will be sent to the tensor core read request scheduler together with the read request for the tensor core;
[0052] The tensor core read request scheduler is mainly used to receive each read request from the tensor core read request queue and the corresponding life cycle value of each read request. Assuming that the read request at the exit of the tensor core read request queue is called the pending execution read request, then, as the life cycle value of the pending execution read request decreases from the initial value, when the life cycle value of the pending execution read request decreases to a preset value (i.e., life cnt = preset value, and the preset value is less than the initial value of the life cycle value), it means that the pending execution read request must be processed in the current clock cycle. That is, the tensor core read request scheduler will issue the pending execution read request to the register file to read the corresponding tensor core operand from the register file, and temporarily store the read tensor core operand in the tensor core operand cache; among them, the priority of the general computing core occupying the register file is only lower than the priority of the tensor core occupying the register file when the life cnt of the pending execution read request = preset value, and before the life cnt of the pending execution read request = preset value, the priority of the general computing core occupying the register file is higher than the priority of the tensor core occupying the register file. That is to say, in the current clock cycle, if the life cnt of the pending execution read request > preset value, the read request of the general computing core is processed first, and the general computing core scheduling unit can issue an instruction to read the register file to occupy the register file for reading operations. If the life cnt of the pending execution read request = preset value, the tensor core read request scheduler must process the pending execution read request. At this time, the general computing core cannot issue an instruction to read the register file, that is, it cannot occupy the register file for reading operations;
[0053] The clock cycle counter is mainly used to record the clock cycle and provide a timing basis for the triggering of the tensor core operand cache. As the clock cycle counter counts, when the count value of the clock cycle counter reaches the preset trigger value, a pop signal will be sent to the tensor core operand cache to trigger the tensor core operand cache to send the tensor core operand to the tensor core execution unit; for example, as Figure 1 shown, when the count value of the clock cycle counter reaches the preset trigger value, a pop signal (i.e., "pop" signal) will be sent to the tensor core operand cache, which is equivalent to a trigger signal. As long as the tensor core operand cache receives a "pop" signal, it will actively send a tensor core operand to the tensor core execution unit. Correspondingly, after the tensor core execution unit finishes the calculation task, it will store the obtained calculation result in the register file.
[0054] It should be noted that after the pending execution read request at the exit of the tensor core read request queue is processed, it is removed from the tensor core read request queue and automatically becomes invalid according to the setting.
[0055] It should be noted that the implementation device for sharing a register bank between a general computing core and a tensor core provided by the embodiments of the present invention can be implemented using various hardware devices. For example, a clock cycle counter can be implemented using a shift register (for example, the clock cycle counter is formed by connecting N registers in series, and it takes N clock cycles to pass from the input of the first register to the output of the last register). The tensor core read request queue and the tensor core operand cache can be implemented based on a memory, register, etc. with storage functions. The embodiments of the present invention do not make specific limitations.
[0056] Exemplarily, the tensor core read request queue can be implemented using a FIFO (First In First Out) queue, and the tensor core operand cache can be implemented using a FIFO queue.
[0057] It should be noted that as Figure 1 shown, the general computing core mainly includes a general computing core instruction fetch unit, a general computing core decoding unit, a general computing core scheduling unit, and a general computing core execution unit. Among them, the general computing core instruction fetch unit is used to fetch the instructions of the general computing core and send them to the general computing core decoding unit; the general computing core decoding unit is used to parse the instructions sent by the general computing core instruction fetch unit, parse out information such as the operation code and operands of the instructions, determine the operations to be performed by the instructions and the relevant information such as the required operands, and send this information to the general computing core scheduling unit for executing the instructions; the general computing core scheduling unit is used to reasonably arrange the execution order of the instructions and the allocation of hardware resources according to the information sent by the general computing core decoding unit, and when the instructions are executed, issue instructions to read the register bank to read the corresponding general computing core operands from the register bank, and send the read general computing core operands to the general computing core execution unit; the general computing core execution unit is used to execute the operations specified by the instructions according to the received general computing core operands and store the operation results in the register bank.
[0058] The implementation device for sharing a register bank between a general computing core and a tensor core provided by the embodiments of the present invention, by introducing a tensor core read request queue and a tensor core operand cache in the tensor core, and cooperating with a tensor core read request scheduler that schedules read requests based on the life cycle value of the tensor core read request, and then collaborating with the general computing core to complete the read bandwidth of the shared register bank between the general computing core and the tensor core, can give a higher priority to the general computing core without blocking the read requests of the tensor core when the general computing core and the tensor core need to process computing tasks simultaneously, maximize the utilization of the read bandwidth of the register bank, thereby effectively improving the overall performance. In an ideal situation, the general computing core and the tensor core can run at full load simultaneously.
[0059] In addition, for the case where the general-purpose computing core and the tensor core share a register bank, in some existing implementations, the general-purpose computing core uses an operand collector to first read the operands and then emit them to the execution unit of the general-purpose computing core for execution. This will result in: on the one hand, when the general-purpose computing core performs operand collection, in fact, the tensor core has a higher priority. Even if there is only one operand required by the general-purpose computing core, it has to wait until the tensor core finishes reading the operands before it can start. During this period, the general-purpose computing core will be idle, thus causing performance waste; on the other hand, when the general-purpose computing core performs operand collection, it needs to save additional information such as instructions and needs to solve data dependency hazards brought by dynamically scheduled instructions, such as RAW (Read After Write), WAW (Write After Write), etc., thereby incurring additional chip area overhead.
[0060] However, an implementation device for sharing a register bank between a general-purpose computing core and a tensor core provided by an embodiment of the present invention can transfer the collection of operands from the existing implementation by the general-purpose computing core to the tensor core for operand caching, so as to achieve an optimal balance between performance and chip area overhead.
[0061] In one optional embodiment, the lifecycle value is set with an initial value, the preset trigger value is the count value corresponding to when the clock cycle counter accumulatively counts up to N clock cycles, and the cumulative count value of the lifecycle value decreasing from the initial value to the preset value does not exceed N.
[0062] Specifically, in combination with the above embodiment, for the lifecycle value of the read request of the tensor core, the initial value of the lifecycle value can be set (denoted as M, and M > preset value). Assume that the lifecycle value starts from the initial value M and automatically decreases by 1 every 1 clock cycle, and the preset value corresponding to the lifecycle value is 1. Then, for each read request of the tensor core, there are M clock cycles available to flexibly schedule this tensor core read request; for the preset trigger value corresponding to the clock cycle counter, the preset trigger value can be set as: the count value corresponding to when the clock cycle counter accumulatively counts up to N clock cycles each time. That is to say, as the clock cycle counter counts, there are multiple preset trigger values and their values are not fixed, but for each read request of the tensor core, the time interval from the start of counting to the count value reaching the preset trigger value of the clock cycle counter is N clock cycles.
[0063] It should be noted that the preset trigger value can be set according to the preparation time (tensor core read operand latency) of the tensor core execution unit before receiving the tensor core operand. For example, if the tensor core execution unit needs N clock cycles to receive the tensor core operand before starting the calculation, the preset trigger value can be set to the count value corresponding to when the clock cycle counter accumulates N clock cycles. As long as the clock cycle counter starts counting from the trigger, when the accumulated count reaches N clock cycles, a tensor core operand is read from the tensor core operand cache and given to the tensor core execution unit.
[0064] Exemplarily, assume that when a read request of a certain tensor core is issued and stored in the tensor core read request queue, the clock cycle counter is triggered to start counting from 1. Every 1 clock cycle, the count value of the clock cycle counter automatically increments by 1. When the accumulated count reaches N clock cycles, the count value of the clock cycle counter is N + 1. Then N + 1 is the preset trigger value of the clock cycle counter corresponding to the read request of this tensor core. That is, when the count value of the clock cycle counter reaches N + 1, it will trigger the tensor core operand cache to send the tensor core operand read when accessing the register file for this tensor core read request to the tensor core execution unit.
[0065] Furthermore, the read requests of the tensor core can be issued continuously. Assume that the read requests A, B, C, D of the tensor core are issued and stored in the tensor core read request queue at the 1st, 2nd, 3rd, and 4th clock cycles respectively. At the same time, the clock cycle counter is triggered to start counting from 1, 2, 3, 4 respectively. Every 1 clock cycle, the count value of the clock cycle counter automatically increments by 1. Assume that the tensor core operands read by A, B, C, D when accessing the register file are a, b, c, d respectively, and a, b, c, d are temporarily stored in the tensor core operand cache. Then, the preset trigger values of the clock cycle counters corresponding to A, B, C, D are N + 1, N + 2, N + 3, N + 4 respectively. That is to say, the stack-out times of a, b, c, d from the tensor core operand cache are N + 1, N + 2, N + 3, N + 4 respectively. That is, when the count values of the clock cycle counter reach N + 1, N + 2, N + 3, N + 4 in sequence, stack-out signals will be continuously sent to the tensor core operand cache to trigger the tensor core operand cache to continuously output the tensor core operands a, b, c, d to the tensor core execution unit.
[0066] It should be noted that since the clock cycle counter is mainly used to generate the stack-out signal to trigger the tensor core operand cache, to ensure that after the tensor core read request is issued, after N clock cycles, the tensor core execution unit can definitely obtain the tensor core operand. Therefore, it is necessary to satisfy that the accumulated count value of the life cycle counter whose life cycle value decreases from the initial value to the preset value does not exceed N.
[0067] Exemplarily, assume that the lifecycle value starts from the initial value M, automatically decrements by 1 every clock cycle, and the preset value corresponding to the lifecycle value is 1, and the clock cycle counter automatically increments by 1 every clock cycle. Then, it is necessary to satisfy 0 < M ≤ N.
[0068] The implementation device for sharing a register bank between a general computing core and a tensor core provided by the embodiments of the present invention can flexibly adjust the values of M and N according to actual needs, thereby improving the flexibility and expandability of the embodiments of the present invention.
[0069] In one optional embodiment, the tensor core read request scheduler is further configured to:
[0070] Send the lifecycle value of the to-be-executed read request to the general computing core scheduling unit; wherein, the general computing core scheduling unit is configured to process the read request of the general computing core according to the lifecycle value of the to-be-executed read request.
[0071] Specifically, in combination with the above embodiments, during the process of sharing the register bank between the general computing core and the tensor core, the tensor core read request scheduler is further configured to send the lifecycle value of the to-be-executed read request in the tensor core read request queue to the general computing core scheduling unit. As Figure 1 shown, the tensor core read request scheduler sends a read request lifecycle (i.e., the "life cnt" signal) to the general computing core scheduling unit, and this "life cnt" signal is the lifecycle value of the to-be-executed read request, which can be used to indicate whether the register bank is occupied by the tensor core. Correspondingly, the general computing core scheduling unit can determine whether the register bank is occupied by the tensor core according to the received lifecycle value of the to-be-executed read request (i.e., the received "life cnt" signal), and process the read request of the general computing core according to the determination result.
[0072] In one optional embodiment, the general computing core scheduling unit processes the read request of the general computing core according to the lifecycle value of the to-be-executed read request, specifically including:
[0073] Determine whether a read request of the general computing core can be issued to the register bank according to the lifecycle value of the to-be-executed read request;
[0074] When the lifecycle value of the to-be-executed read request is not the preset value, it is determined that a read request of the general computing core can be issued to the register bank;
[0075] When the lifecycle value of the to-be-executed read request is the preset value, it is determined that a read request of the general computing core cannot be issued to the register bank.
[0076] Specifically, in combination with the above embodiments, when the general computing core scheduling unit processes the read request of the general computing core according to the life cycle value of the to-be-executed read request received (i.e., the "life cnt" signal received), it can determine whether the register bank is occupied by the tensor core according to the life cycle value of the to-be-executed read request received, so as to determine whether it can issue the read request of the general computing core to the register bank; if the life cycle value of the to-be-executed read request is not the preset value, it indicates that the register bank is not occupied by the tensor core, then it is determined that the read request of the general computing core can be processed, that is, the read request of the general computing core can be issued to the register bank to read the general computing core operands from the register bank; if the life cycle value of the to-be-executed read request is the preset value, it indicates that the register bank has been occupied by the tensor core, then it is determined that the read request of the general computing core cannot be processed, that is, the read request of the general computing core cannot be issued to the register bank.
[0077] Exemplarily, assume that the initial value M of the life cycle value is 4, it automatically decrements by 1 every 1 clock cycle, and the preset value corresponding to the life cycle value is 1. Then, in the current clock cycle, if the "life cnt" signal indicates that life cnt > 1 (i.e., life cnt ≠ 1, life cnt = 4 or 3 or 2), the general computing core scheduling unit can occupy the register bank and preferentially process the read request of the general computing core; if the "life cnt" signal indicates that life cnt = 1, the general computing core scheduling unit cannot occupy the register bank to process the read request of the general computing core. At this time, the register bank is occupied by the tensor core to process the to-be-executed read request with life cnt = 1 at the exit of the tensor core read request queue.
[0078] An implementation device for sharing a register bank between a general computing core and a tensor core provided by an embodiment of the present invention can, by setting the life cycle of the read request of the tensor core and alternately occupying the read bandwidth of the register bank in each clock cycle (taking one clock cycle as a unit) to respectively execute the read request of the general computing core and the read request of the tensor core, give the idle time of the tensor core to the general computing core for use, reduce the idle time of the general computing core, thereby maximizing the utilization of the read bandwidth of the register bank and effectively improving the overall performance.
[0079] In one optional embodiment, the general computing core scheduling unit is further configured to:
[0080] When it is determined that the read request of the general computing core can be issued to the register bank, determine the maximum number of general computing core operands that are currently allowed to be read in the read request of the general computing core according to the current value of the life cycle value of the to-be-executed read request.
[0081] Specifically, in combination with the above embodiments, when the general computing core scheduling unit determines that it can issue a read request of the general computing core to the register bank according to the life cycle value of the to-be-executed read request received (i.e., the received "life cnt" signal), it can further determine, according to the life cycle value of the to-be-executed read request received, the maximum number of general computing core operands allowed to be read from the register bank in a read request currently issued by the general computing core scheduling unit (a read request can read one or more operands). That is, according to different values of the "life cnt" signal, the general computing core scheduling unit can issue instructions for reading different numbers of general computing core operands.
[0082] Exemplarily, assume that the initial value M of the life cycle value is 4, it automatically decrements by 1 every 1 clock cycle, and the preset value corresponding to the life cycle value is 1. Then, in the current clock cycle, if the "life cnt" signal indicates that life cnt = 4, the general computing core scheduling unit can continuously read for at most 3 clock cycles and can issue instructions for at most 3 general computing core operands at most; if the "life cnt" signal indicates that life cnt = 3, the general computing core scheduling unit can continuously read for at most 2 clock cycles and can issue instructions for at most 2 general computing core operands at most; if the "life cnt" signal indicates that life cnt = 2, the general computing core scheduling unit can continuously read for at most 1 clock cycle and can issue instructions for at most 1 general computing core operand at most.
[0083] It can be understood that if the "life cnt" signal indicates that life cnt = 1, the general computing core scheduling unit cannot read general computing core operands from the register bank, and at this time, it can issue instructions that do not require reading general computing core operands.
[0084] In one optional embodiment, the tensor core read request scheduler is further configured to:
[0085] Receive the read operand information sent by the general computing core scheduling unit;
[0086] Process the read requests in the tensor core read request queue according to the value of the read operand information; wherein, the value of the read operand information is used to indicate whether the register bank is occupied by the general computing core.
[0087] Specifically, in combination with the above embodiments, during the process of sharing the register bank between the general computing core and the tensor core, the tensor core read request scheduler is further configured to receive the read operand information sent by the general computing core scheduling unit, such as Figure 1As shown, the general computing core scheduling unit sends read operand information (i.e., the "src" signal) to the tensor core read request scheduler. The value of this "src" signal can be used to indicate whether the register file is occupied by the general computing core. Correspondingly, the tensor core read request scheduler can determine whether the register file is occupied by the general computing core based on the value of the received read operand information (i.e., the received "src" signal), and process the read requests of the tensor cores in the tensor core read request queue according to the determination result.
[0088] In one optional embodiment, the tensor core read request scheduler processes the read requests in the tensor core read request queue according to the value of the read operand information, specifically including:
[0089] Determine whether the read requests in the tensor core read request queue can be issued to the register file according to the value of the read operand information;
[0090] When the value of the read operand information is the first preset value, determine whether the life cycle value of the to-be-executed read request is the preset value; wherein, the first preset value is used to indicate that the register file has been occupied by the general computing core;
[0091] If so, it is determined that the read requests in the tensor core read request queue can be issued to the register file;
[0092] If not, it is determined that the read requests in the tensor core read request queue cannot be issued to the register file.
[0093] Specifically, in combination with the above embodiment, when the tensor core read request scheduler processes the read requests of the tensor cores in the tensor core read request queue according to the value of the received read operand information (i.e., the received "src" signal), it can determine whether the register file is occupied by the general computing core according to the value of the received read operand information to determine whether the read requests of the tensor cores in the tensor core read request queue can be issued to the register file; if the value of the read operand information is the first preset value, and the first preset value is used to indicate that the register file has been occupied by the general computing core, then it can further determine whether the life cycle value of the to-be-executed read request in the tensor core read request queue is the preset value; if the life cycle value of the to-be-executed read request is the preset value, at this time, although the register file has been occupied by the general computing core, the tensor core has a higher priority to occupy the register file, then it is determined that the read requests of the tensor cores in the tensor core read request queue can be issued to the register file; if the life cycle value of the to-be-executed read request is not the preset value, at this time, the register file has been occupied by the general computing core, and the general computing core has a higher priority to occupy the register file, then it is determined that the read requests of the tensor cores in the tensor core read request queue cannot be issued to the register file.
[0094] Exemplarily, assume that the initial value M of the lifecycle value is 4, and it automatically decrements by 1 every clock cycle. Also, assume that the preset value corresponding to the lifecycle value is 1, and assume that the first preset value ≠ 0. That is, when the "src" signal indicates that src ≠ 0, it means that the register bank is being occupied by the general computing core. Then, in the current clock cycle, if the tensor core read request scheduler receives src ≠ 0, the tensor core read request scheduler will further determine whether the lifecycle value of the pending read request in the tensor core read request queue satisfies life cnt = 1; if life cnt = 1, the tensor core read request scheduler can issue a read request for the tensor core in the tensor core read request queue to the register bank; if life cnt ≠ 1, that is, life cnt = 4 or 3 or 2, the tensor core read request scheduler cannot issue a read request for the tensor core in the tensor core read request queue to the register bank; that is to say, on the premise that src ≠ 0, only when life cnt = 1 can the tensor core read request scheduler issue a read request for the tensor core in the tensor core read request queue to the register bank.
[0095] In one alternative embodiment, when the tensor core read request scheduler processes the read requests in the tensor core read request queue according to the value of the read operand information, it further includes:
[0096] When the value of the read operand information is the second preset value, it is determined that a read request for the tensor core in the tensor core read request queue can be issued to the register bank; wherein, the second preset value is used to indicate that the register bank is not occupied by the general computing core.
[0097] Specifically, in combination with the above embodiment, when the tensor core read request scheduler determines whether the register bank is occupied by the general computing core according to the value of the received read operand information (i.e., the received "src" signal) to determine whether a read request for the tensor core in the tensor core read request queue can be issued to the register bank, if the value of the read operand information is the second preset value, and the second preset value is used to indicate that the register bank is not occupied by the general computing core, it is determined that a read request for the tensor core in the tensor core read request queue can be issued to the register bank.
[0098] Exemplarily, assume that the initial value M of the lifecycle value is 4, and it automatically decreases by 1 every clock cycle. Also, assume that the preset value corresponding to the lifecycle value is 1, and assume that the second preset value = 0. That is, when the "src" signal indicates src = 0, it means that the register bank is not occupied by the general computing core. Then, in the current clock cycle, if the tensor core read request scheduler receives src = 0, regardless of whether life cnt = 1 is satisfied, the tensor core read request scheduler can issue a read request for the tensor core in the tensor core read request queue to the register bank; that is to say, on the premise that src = 0, whether life cnt = 4 or 3 or 2 or 1, the tensor core read request scheduler can issue a read request for the tensor core in the tensor core read request queue to the register bank.
[0099] In one alternative embodiment, the tensor core read request scheduler is further configured to:
[0100] When it is determined that a read request in the tensor core read request queue can be issued to the register bank, issue the read request in the tensor core read request queue to the register bank to read tensor core operands from the register bank and store them in the tensor core operand cache.
[0101] Specifically, in combination with the above embodiment, when the tensor core read request scheduler determines that a read request for the tensor core in the tensor core read request queue can be issued to the register bank according to the value of the received read operand information (i.e., the received "src" signal), it can issue a read request for the tensor core in the tensor core read request queue to the register bank according to the arrangement order of the read requests for the tensor core in the tensor core read request queue (which may include read requests for tensor cores with life cnt = 1 and read requests for tensor cores with life cnt ≠ 1) to read the corresponding tensor core operands from the register bank and temporarily store them in the tensor core operand cache. After that, it can be understood that when the count value of the clock cycle counter reaches the preset trigger value, a stack-out signal (i.e., the "pop" signal) will be sent to the tensor core operand cache to trigger the tensor core operand cache to actively send a tensor core operand to the tensor core execution unit.
[0102] To solve the above problems, an embodiment of the present invention further provides a method for implementing a shared register bank between a general computing core and a tensor core. Refer to Figure 2 As shown, it is a schematic flowchart of a preferred embodiment of a method for implementing a shared register bank between a general computing core and a tensor core provided by the present invention. The implementation method is applicable to the implementation device described in any of the above embodiments. The implementation method includes steps S11 to S14:
[0103] Step S11: Issue a read request through the tensor core read request unit and store it in the tensor core read request queue. At the same time, trigger the lifecycle counter and the clock cycle counter to start counting;
[0104] Step S12: Record the lifecycle value of each read request in the tensor core read request queue through a lifecycle counter, and the lifecycle value decreases as the lifecycle counter counts.
[0105] Step S13: Receive the read request and its lifecycle value from the tensor core read request queue through a tensor core read request scheduler, and when the lifecycle value of the pending read request reaches a preset value, transmit the pending read request to the register bank to read tensor core operands from the register bank and store them in the tensor core operand cache; wherein, the pending read request is the read request at the exit of the tensor core read request queue; the priority of the general computing core occupying the register bank is higher than that of the tensor core before the lifecycle value of the pending read request reaches the preset value.
[0106] Step S14: When the count value of the clock cycle counter reaches a preset trigger value, send a stack-out signal to the tensor core operand cache to trigger the tensor core operand cache to send tensor core operands to the tensor core execution unit.
[0107] Preferably, the lifecycle value is set with an initial value, the preset trigger value is the count value corresponding to when the clock cycle counter accumulatively counts to N clock cycles, and the cumulative count value of the lifecycle value decreasing from the initial value to the preset value does not exceed N.
[0108] Preferably, the implementation method further includes:
[0109] Send the lifecycle value of the pending read request to the general computing core scheduler through the tensor core read request scheduler; wherein, the general computing core scheduler is used to process the read request of the general computing core according to the lifecycle value of the pending read request.
[0110] Preferably, the general computing core scheduler processes the read request of the general computing core according to the lifecycle value of the pending read request, specifically including:
[0111] Judge whether a read request of the general computing core can be transmitted to the register bank according to the lifecycle value of the pending read request;
[0112] When the lifecycle value of the pending read request is not the preset value, it is determined that a read request of the general computing core can be transmitted to the register bank;
[0113] When the lifecycle value of the pending read request is the preset value, it is determined that a read request of the general computing core cannot be transmitted to the register bank.
[0114] Preferably, the general computing core scheduler is further used for:
[0115] When determining that a read request for a general computing core can be issued to the register bank, according to the current value of the life cycle value of the to-be-executed read request, determine the maximum number of general computing core operands that can be read currently in the read request of the general computing core.
[0116] Preferably, the implementation method further includes:
[0117] Receive the read operand information sent by the general computing core scheduling unit through the tensor core read request scheduler, and process the read requests in the tensor core read request queue according to the value of the read operand information; wherein, the value of the read operand information is used to indicate whether the register bank is occupied by the general computing core.
[0118] Preferably, the processing of the read requests in the tensor core read request queue according to the value of the read operand information specifically includes:
[0119] Judge whether a read request in the tensor core read request queue can be issued to the register bank according to the value of the read operand information;
[0120] When the value of the read operand information is a first preset value, judge whether the life cycle value of the to-be-executed read request is the preset value; wherein, the first preset value is used to indicate that the register bank has been occupied by the general computing core;
[0121] If so, determine that a read request in the tensor core read request queue can be issued to the register bank;
[0122] If not, determine that a read request in the tensor core read request queue cannot be issued to the register bank.
[0123] Preferably, the processing of the read requests in the tensor core read request queue according to the value of the read operand information further includes:
[0124] When the value of the read operand information is a second preset value, determine that a read request in the tensor core read request queue can be issued to the register bank; wherein, the second preset value is used to indicate that the register bank is not occupied by the general computing core.
[0125] Preferably, the implementation method further includes:
[0126] When determining that a read request in the tensor core read request queue can be issued to the register bank, issue the read requests in the tensor core read request queue to the register bank through the tensor core read request scheduler, so as to read tensor core operands from the register bank and store them in the tensor core operand cache.
[0127] It should be noted that the implementation method of sharing a register bank between a general computing core and a tensor core provided by the embodiments of the present invention can implement all the processing flows in the implementation device described in any of the above embodiments. The specific implementation schemes and the achieved technical effects corresponding to the implementation method are respectively the same as those of the implementation device described in the above embodiments, and will not be elaborated here.
[0128] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. An implementation device for a general computing core and a tensor core to share a register bank, characterized in that Applied to tensor cores, including: A tensor core read request unit, which is used to issue a read request and store it in the tensor core read request queue. At the same time, it triggers the life cycle counter and the clock cycle counter to start counting; A life cycle counter, which is used to record the life cycle value of each read request in the tensor core read request queue, and the life cycle value decreases as the life cycle counter counts; A tensor core read request scheduler, which is used to receive the read request and its life cycle value from the tensor core read request queue, and when the life cycle value of the to-be-executed read request reaches a preset value, it emits the to-be-executed read request to the register bank to read tensor core operands from the register bank and store them in the tensor core operand cache; wherein, the to-be-executed read request is the read request at the exit of the tensor core read request queue; the priority of the general computing core occupying the register bank is higher than that of the tensor core before the life cycle value of the to-be-executed read request reaches the preset value; A clock cycle counter, which is used to send a stack-out signal to the tensor core operand cache when the count value reaches a preset trigger value, so as to trigger the tensor core operand cache to send tensor core operands to the tensor core execution unit.
2. The implementation device for sharing a register bank by a general computing core and a tensor core according to claim 1, wherein The life cycle value is set with an initial value, the preset trigger value is the count value corresponding to when the clock cycle counter accumulatively counts to N clock cycles, and the accumulative count value of the life cycle value decreasing from the initial value to the preset value does not exceed N.
3. The implementation device for sharing a register bank between a general computing core and a tensor core according to claim 1, wherein The tensor core read request scheduler is further used for: Sending the life cycle value of the to-be-executed read request to the general computing core scheduling unit; wherein, the general computing core scheduling unit is used to process the read request of the general computing core according to the life cycle value of the to-be-executed read request.
4. The implementation device for sharing a register bank by a general computing core and a tensor core according to claim 3, wherein The general computing core scheduling unit processes the read request of the general computing core according to the life cycle value of the to-be-executed read request, specifically including: Judging whether a read request of the general computing core can be emitted to the register bank according to the life cycle value of the to-be-executed read request; When the life cycle value of the to-be-executed read request is not the preset value, it is determined that a read request of the general computing core can be emitted to the register bank; When the life cycle value of the to-be-executed read request is the preset value, it is determined that a read request of the general computing core cannot be emitted to the register bank.
5. The implementation device for sharing a register bank by a general computing core and a tensor core according to claim 4, characterized in that The general computing core scheduling unit is further used for: When it is determined that a read request of the general computing core can be emitted to the register bank, according to the current value of the life cycle value of the to-be-executed read request, determining the maximum number of general computing core operands that can be read currently in the read request of the general computing core.
6. The implementation device for sharing a register bank by a general computing core and a tensor core according to claim 1, wherein The tensor core read request scheduler is further used for: Receiving the read operand information sent by the general computing core scheduling unit; Processing the read requests in the tensor core read request queue according to the value of the read operand information; wherein, the value of the read operand information is used to indicate whether the register bank is occupied by the general computing core.
7. The implementation device for sharing a register bank between a general computing core and a tensor core according to claim 6, wherein The tensor core read request scheduler processes the read requests in the tensor core read request queue according to the value of the read operand information, specifically including: Determine whether a read request in the tensor core read request queue can be issued to the register bank according to the value of the read operand information; When the value of the read operand information is a first preset value, determine whether the life cycle value of the to-be-executed read request is the preset value; wherein, the first preset value is used to indicate that the register bank is occupied by a general computing core; If so, determine that a read request in the tensor core read request queue can be issued to the register bank; If not, determine that a read request in the tensor core read request queue cannot be issued to the register bank.
8. The implementation device for sharing a register bank between a general computing core and a tensor core according to claim 7, wherein The tensor core read request scheduler processes the read requests in the tensor core read request queue according to the value of the read operand information, and further includes: When the value of the read operand information is a second preset value, determine that a read request in the tensor core read request queue can be issued to the register bank; wherein, the second preset value is used to indicate that the register bank is not occupied by a general computing core.
9. The implementation device for sharing a register bank by a general computing core and a tensor core according to claim 7 or 8, wherein The tensor core read request scheduler is further configured to: When it is determined that a read request in the tensor core read request queue can be issued to the register bank, issue the read request in the tensor core read request queue to the register bank to read tensor core operands from the register bank and store them in the tensor core operand cache.
10. A method for implementing a shared register bank between a general computing core and a tensor core, characterized in that, Applicable to the implementation device according to any one of claims 1 to 9, including: Issue a read request through the tensor core read request unit and store it in the tensor core read request queue. At the same time, trigger the life cycle counter and the clock cycle counter to start counting; Record the life cycle value of each read request in the tensor core read request queue through the life cycle counter, and the life cycle value decreases as the life cycle counter counts; Receive a read request and its life cycle value from the tensor core read request queue through the tensor core read request scheduler, and when the life cycle value of the to-be-executed read request reaches the preset value, issue the to-be-executed read request to the register bank to read tensor core operands from the register bank and store them in the tensor core operand cache; wherein, the to-be-executed read request is the read request at the exit of the tensor core read request queue; the priority of the general computing core occupying the register bank is higher than that of the tensor core before the life cycle value of the to-be-executed read request reaches the preset value; When the count value of the clock cycle counter reaches a preset trigger value, send a stack-out signal to the tensor core operand cache to trigger the tensor core operand cache to send tensor core operands to the tensor core execution unit.
Citation Information
Cited By
Artificial intelligence chip and operation method thereof
CN120655494A