Data processing integrated circuit
By designing the computing unit circuit and tensor core circuit in the data processing integrated circuit, the problem of how to efficiently execute multiple thread groups in the same working group is solved, and efficient tensor computing performance is achieved.
Patent Information
- Application Number
- CN202210386429.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-11
AI Technical Summary
How to design hardware to efficiently execute multiple thread groups in the same workgroup, especially when performing tensor calculations.
A data processing integrated circuit is designed, including a computing unit circuit, each computing unit circuit including a tensor core circuit. The tensor kernel circuit is aligned with thread groups, and can perform tensor calculations to achieve efficient processing of multiple thread groups.
With this design, the data processing integrated circuit can efficiently execute multiple thread groups in a working group, improving the efficiency and performance of tensor computing.
Smart Images

Figure CN114780236B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an integrated circuit, and particularly to a data processing integrated circuit. Background Art
[0002] Data processing integrated circuits such as central processing units (CPUs), graphics processing units (GPUs), and general-purpose computing on GPUs (GPGPUs) can execute programs to complete various functions such as convolutional neural network (CNN) operations and artificial intelligence operations. Generally, a program includes multiple workgroups, each workgroup includes multiple warps, and each warp includes multiple threads. Threads in the same workgroup can be grouped according to a scheduling unit and then scheduled to the hardware group by group for execution. This scheduling unit is called a warp. How to design the hardware to execute multiple warps in the same workgroup is one of the many topics in this technical field. Summary of the Invention
[0003] The present invention provides a data processing integrated circuit to execute multiple warps in a workgroup.
[0004] In an embodiment according to the present invention, the data processing integrated circuit includes at least one compute unit circuit. Each compute unit circuit is used to execute multiple warps in a workgroup. Each compute unit circuit includes a single tensor core circuit. The tensor core circuit is used to perform tensor calculations on one warp of the multiple warps.
[0005] Based on the above, the compute unit circuit is aligned with the workgroup, and the tensor core is aligned with the warp. The data processing integrated circuit can execute multiple warps in a workgroup, and the single tensor core circuit configured in the compute unit circuit can perform tensor calculations on the warp. Brief Description of the Drawings
[0006] Figure 1 It is a schematic diagram of a circuit block of a data processing integrated circuit according to an embodiment of the present invention.
[0007] Figure 2 It is a circuit block diagram of a computing unit circuit shown according to an embodiment of the present invention.
[0008] Figure 3 It is a circuit block diagram of a computing unit circuit shown according to another embodiment of the present invention.
[0009] Figure 4 It is a circuit block diagram of a computing unit circuit shown according to still another embodiment of the present invention.
[0010] Description of Reference Numerals
[0011] 100: Data Processing Integrated Circuit
[0012] 110_1, 110_n: Computing Unit Circuit
[0013] 111_1, 111_n: Tensor Core Circuit
[0014] 120: Secondary Cache
[0015] 210, 310, 410: Load Store Unit Circuit
[0016] 220: Data Bus
[0017] 230, 330, 430: Execution Unit Circuit
[0018] 320, 420: Random Access Memory
[0019] 421: Primary Cache
[0020] 422: Workgroup Shared Memory Detailed Description of the Invention
[0021] Reference will now be made in detail to the exemplary embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numerals are used in the drawings and the description to refer to the same or like parts.
[0022] As used throughout the specification (including the claims) of this case, the term "coupled (or connected)" may refer to any direct or indirect connection means. For example, if it is described in the text that a first device is coupled (or connected) to a second device, it should be interpreted that the first device can be directly connected to the second device, or the first device can be indirectly connected to the second device through other devices or certain connection means. The terms "first", "second", etc. mentioned throughout the specification (including the claims) of this case are used to name elements, rather than to limit the upper or lower limit of the number of elements, nor to limit the order of elements. In addition, wherever possible, components / structures / steps with the same reference numerals in the drawings and embodiments represent the same or similar parts. Components / structures / steps with the same reference numerals or the same terms used in different embodiments can be referred to the relevant descriptions.
[0023] Figure 1 FIG. is a schematic diagram of a circuit block of a data processing integrated circuit 100 according to an embodiment of the present invention. According to an actual implementation, the data processing integrated circuit 100 may include a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose computing on GPU (GPGPU), a microcontroller, a microprocessor, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), and / or other data processing integrated circuits. The data processing integrated circuit 100 can execute a program to complete various functions such as convolutional neural network (CNN) operations and artificial intelligence operations.
[0024] The data processing integrated circuit 100 includes at least one computing unit (CU) circuit, such as Figure 1The n computing unit circuits 110_1, …, 110_n are shown. The number n of the computing unit circuits 110_1 to 110_n can be determined according to the actual design. For example, in some embodiments, the number n can be 1, 2 or other integers. The computing unit circuits are aligned with a workgroup (WG), that is, one computing unit circuit can execute one workgroup in a program in one batch.
[0025] Each of the computing unit circuits 110_1 to 110_n can execute multiple warps in a corresponding workgroup. In Figure 1 the illustrated embodiment, the tensor core (TC) circuits 111_1 to 111_n are aligned with the computing unit circuits 110_1 to 110_n. That is, each of the computing unit circuits 110_1 to 110_n includes a single tensor core circuit (i.e., only includes one tensor core circuit). For example, the computing unit circuit 110_1 includes a single tensor core circuit 111_1, and the computing unit circuit 110_n includes a single tensor core circuit 111_n. The tensor core is aligned with a warp, that is, one tensor core circuit can execute one warp in the same workgroup in one batch. For example, the tensor core circuit 111_1 can perform tensor calculations on a warp. In the same computing unit circuit, the synchronization scope between the tensor core circuit and other processing elements is the workgroup. In addition, the synchronization scope between multiple warps in the same workgroup can also be the workgroup. After synchronization, the calculation results of the tensor core can be used by multiple warps in the same workgroup. For example, each warp in multiple warps can use different parts of the calculation results of the tensor core.
[0026] In summary, each of the computing unit circuits 110_1 to 110_n can execute multiple warps in a corresponding workgroup. The tensor core circuits 111_1 to 111_n are aligned with the computing unit circuits 110_1 to 110_n. The tensor core circuits 111_1 to 111_n can perform tensor calculations on warps in different workgroups. The specific implementation manners of the computing unit circuits 110_1 to 110_n are not limited in this embodiment. Multiple examples will be introduced for the computing unit circuits 110_1 to 110_n below.
[0027] Figure 2 is a schematic block diagram of the computing unit circuit 110_1, shown according to an embodiment of the present invention. In Figure 2In the illustrated embodiment, the data processing integrated circuit 100 further includes a level 2 cache (L2 cache) 120. The level 2 cache 120 is coupled to the computing unit circuits 110_1 to 110_n. The computing unit circuits 110_1 to 110_n can access the main memory (not shown) outside the data processing integrated circuit 100 through the level 2 cache 120.
[0028] In Figure 2 In the illustrated embodiment, the computing unit circuit 110_1 includes a load / store unit (LSU) circuit 210, a data bus 220, a tensor core circuit 111_1, and at least one execution unit (EU) circuit 230. The number of execution unit circuits 230 can be determined according to the actual design. Each execution unit circuit 230 can perform vector calculations on a warp (in multiple thread groups) in one batch (iteration). Other computing unit circuits (such as the computing unit circuit 110_n) can refer to the relevant description of the computing unit circuit 110_1 and be analogized, so they will not be elaborated here.
[0029] In the computing unit circuit 110_1, the load / store unit circuit 210 is coupled to the data bus 220 and the level 2 cache 120. The load / store unit circuit 210 can access the main memory (not shown) outside the data processing integrated circuit 100 through the level 2 cache 120. The tensor core circuit 111_1 and the execution unit circuit 230 are coupled to the data bus 220. The tensor core circuit 111_1 can access the level 2 cache 120 through the data bus 220 and the load / store unit circuit 210.
[0030] The tensor core circuit 111_1 can provide the calculation result of the tensor calculation to the execution unit circuit 230 in any way. For example, in some embodiments, the tensor core circuit 111_1 can provide the calculation result of the tensor calculation to the register file in the execution unit circuit 230 through the data bus 210 for use by the execution unit circuit 230. Based on the actual application, the calculation result of the tensor calculation can include the General matrix multiply (GEMM) result and / or other calculation results.
[0031] In some other embodiments, the tensor core circuit 111_1 can store the calculation results of tensor calculations in the workgroup shared memory (GSM) in the main memory (not shown) outside the data processing integrated circuit 100 through the data bus 220, the load / store unit circuit 210, and the secondary cache 120. The execution unit circuit 230 can retrieve the calculation results provided by the tensor core circuit 111_1 from the workgroup shared memory (not shown) through the data bus 220, the load / store unit circuit 210, and the secondary cache 120.
[0032] Figure 3 FIG. 4 is a schematic block diagram of the computing unit circuit 110_1 according to another embodiment of the present invention. Figure 3 The illustrated data processing integrated circuit 100 further includes a secondary cache (L2 cache) 120. Figure 3 The illustrated secondary cache 120 and the computing unit circuits 110_1 to 110_n can be referred to Figure 2 the relevant descriptions of the illustrated secondary cache 120 and the computing unit circuits 110_1 to 110_n and analogized, so details are not described herein again. In Figure 3 the illustrated embodiment, the computing unit circuit 110_1 includes a load / store unit (LSU) circuit 310, a random access memory (RAM) 320, a tensor core circuit 111_1, and at least one execution unit (EU) circuit 330. Figure 3 The illustrated load / store unit circuit 310, tensor core circuit 111_1, and execution unit circuit 330 can be referred to Figure 2 the relevant descriptions of the illustrated load / store unit circuit 210, tensor core circuit 111_1, and execution unit circuit 230 and analogized, so details are not described herein again.
[0033] Other computing unit circuits (such as the computing unit circuit 110_n) can be referred to the relevant descriptions of the computing unit circuit 110_1 and analogized, so details are not described herein again. In Figure 3In the illustrated computing unit circuit 110_1, the load / store unit circuit 310 is coupled to the random access memory 320 and the secondary cache 120. The load / store unit circuit 310 can load data from a main memory (not shown) external to the data processing integrated circuit 100 into the random access memory 320 through the secondary cache 120. The load / store unit circuit 310 can also store the data in the random access memory 320 into the main memory through the secondary cache 120. The tensor core circuit 111_1 and the execution unit circuit 330 are coupled to the random access memory 320. The tensor core circuit 111_1 can store the calculation result of the tensor calculation in the random access memory 320 for use by the execution unit circuit 330.
[0034] In some practical application examples, the random access memory 320 may include a level 1 cache (L1 cache). The tensor core circuit 111_1 can store the calculation result of the tensor calculation in the level 1 cache for use by the execution unit circuit 330. Alternatively, the tensor core circuit 111_1 can store the calculation result of the tensor calculation in a workgroup shared memory (GSM) in the main memory (not shown) external to the data processing integrated circuit 100 through the level 1 cache, the load / store unit circuit 310, and the secondary cache 120. The execution unit circuit 330 can retrieve the calculation result provided by the tensor core circuit 111_1 from the workgroup shared memory (not shown) through the level 1 cache, the load / store unit circuit 310, and the secondary cache 120.
[0035] Figure 4 It is a schematic block diagram of the computing unit circuit 110_1 shown in accordance with another embodiment of the present invention. Figure 4 The illustrated data processing integrated circuit 100 further includes a secondary cache (L2 cache) 120. Figure 4 For the illustrated secondary cache 120 and the computing unit circuits 110_1 to 110_n, reference can be made to Figure 2 the relevant descriptions of the illustrated secondary cache 120 and the computing unit circuits 110_1 to 110_n and analogized, so details are not repeated here. In Figure 4 the illustrated embodiment, the computing unit circuit 110_1 includes a load / store unit (LSU) circuit 410, a random access memory (RAM) 420, a tensor core circuit 111_1, and at least one execution unit (EU) circuit 430. Other computing unit circuits (such as the computing unit circuit 110_n) can be analogized with reference to the relevant descriptions of the computing unit circuit 110_1, so details are not repeated here.
[0036] Figure 4 The illustrated load / store unit circuit 410 is coupled to the random access memory 420 and the secondary cache 120. Figure 4The illustrated load / store unit circuit 410, random access memory 420, tensor core circuit 111_1, and execution unit circuit 430 can be referred to Figure 3 the relevant descriptions of the illustrated load / store unit circuit 310, random access memory 320, tensor core circuit 111_1, and execution unit circuit 330 and be analogized accordingly, and thus will not be elaborated here. In Figure 4 the illustrated computing unit circuit 110_1, the random access memory 420 can include a level-1 cache (L1 cache) 421 and a workgroup shared memory (GSM) 422. The tensor core circuit 111_1 can store the calculation results of tensor calculations in the workgroup shared memory 422 for use by the execution unit circuit 430. For example, a general matrix multiply (GEMM) instruction can provide a GSM address pointing to the workgroup shared memory 422. The tensor core circuit 111_1 can execute the GEMM instruction and output the GEMM result (the calculation result of tensor calculations) to the workgroup shared memory 422 according to the GSM address provided by the GEMM instruction.
[0037] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data processing integrated circuit, characterized in that, The data processing integrated circuit includes: At least one computing unit circuit, each of the computing unit circuits being configured to execute multiple thread groups in a workgroup, wherein each of the at least one computing unit circuits includes a single tensor core circuit, and the tensor core circuit is configured to perform tensor calculations on one of the multiple thread groups; A secondary cache, coupled to the at least one computing unit circuit, wherein any one of the at least one computing unit circuits includes: A data bus, wherein the tensor core circuit is coupled to the data bus; A load / store unit circuit, coupled to the data bus and the secondary cache, wherein the tensor core circuit accesses the secondary cache through the data bus and the load / store unit circuit; and At least one execution unit circuit, coupled to the data bus, each of the execution unit circuits being configured to perform vector calculations on one of the multiple thread groups, wherein the tensor core circuit stores the calculation results of the tensor calculations in the workgroup shared memory in the main memory outside the data processing integrated circuit through the data bus, the load / store unit circuit, and the secondary cache, and the at least one execution unit circuit retrieves the calculation results from the workgroup shared memory through the data bus, the load / store unit circuit, and the secondary cache.
2. The data processing integrated circuit according to claim 1, wherein the tensor core circuit further provides the calculation results of the tensor calculations to a register file in the at least one execution unit circuit for use by the at least one execution unit circuit through the data bus.
3. The data processing integrated circuit according to claim 1, wherein The synchronization scope within each of the computing unit circuits and / or between the multiple thread groups in each workgroup is the workgroup.
4. A data processing integrated circuit, characterized in that, The data processing integrated circuit includes: At least one computing unit circuit, each of the computing unit circuits being configured to execute multiple thread groups in a workgroup, wherein each of the at least one computing unit circuits includes a single tensor core circuit, and the tensor core circuit is configured to perform tensor calculations on one of the multiple thread groups; A secondary cache, coupled to the at least one computing unit circuit, wherein any one of the at least one computing unit circuits includes: A random access memory, wherein the tensor core circuit is directly coupled to the random access memory; and A load / store unit circuit, directly coupled to the random access memory and the secondary cache, wherein the load / store unit circuit is configured to load data from the main memory outside the data processing integrated circuit to the random access memory through the secondary cache, and the load / store unit circuit is configured to store the data in the random access memory to the main memory through the secondary cache.
5. The data processing integrated circuit according to claim 4, characterized in that, Any one of the at least one computing unit circuits further includes: At least one execution unit circuit, coupled to the random access memory, each of the execution unit circuits being configured to perform vector calculations on one of the multiple thread groups, The tensor core circuit stores the calculation result of the tensor calculation in the random access memory for use by the at least one execution unit circuit.
6. The data processing integrated circuit according to claim 4, wherein Any one of the at least one computing unit circuit further includes: At least one execution unit circuit, coupled to the random access memory, each execution unit circuit is used to perform vector calculations on one of the multiple thread groups. Wherein the random access memory includes a level-1 cache. The tensor core circuit stores the calculation result of the tensor calculation in the level-1 cache for use by the at least one execution unit circuit.
7. The data processing integrated circuit according to claim 4, wherein Any one of the at least one computing unit circuit further includes: At least one execution unit circuit, coupled to the random access memory, each execution unit circuit is used to perform vector calculations on one of the multiple thread groups. Wherein the random access memory includes a level-1 cache. The tensor core circuit stores the calculation result of the tensor calculation in the workgroup shared memory in the main memory outside the data processing integrated circuit through the level-1 cache, the load / store unit circuit, and the level-2 cache, and The at least one execution unit circuit retrieves the calculation result from the workgroup shared memory through the level-1 cache, the load / store unit circuit, and the level-2 cache.
8. The data processing integrated circuit according to claim 4, wherein Any one of the at least one computing unit circuit further includes: At least one execution unit circuit, coupled to the random access memory, each execution unit circuit is used to perform vector calculations on one of the multiple thread groups. Wherein the random access memory includes a level-1 cache and a workgroup shared memory. The tensor core circuit stores the calculation result of the tensor calculation in the workgroup shared memory for use by the at least one execution unit circuit.
9. The data processing integrated circuit according to claim 4, wherein The synchronization range within each computing unit circuit and / or between multiple thread groups in each workgroup is the workgroup.
Citation Information
Patent Citations
Techniques for orchestrating stages of thread synchronization
CN113495761A