Computing device and its operating method

By introducing data memory access units and global cache into the computing device of AI chips, the parallel execution of tensor kernels and vector kernels is achieved, solving the problem of underutilization of computing resources and improving computing efficiency.

CN118394525BActive Publication Date: 2026-04-24SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI BIREN TECH CO LTD
Filing Date
2024-05-22
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing technologies, different types of computing cores in the execution unit of AI chips cannot be executed in parallel, resulting in underutilization of computing resources.

Method used

By introducing a data memory access unit into the computing device, data transfer between the global cache and the shared cache is realized, and different types of computing cores, such as tensor cores and vector cores, are executed in parallel to form thread groups, so that different types of computing cores can run in parallel in the same execution unit.

Benefits of technology

It achieves full utilization of computing resources and improves computing efficiency and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118394525B_ABST
    Figure CN118394525B_ABST
Patent Text Reader

Abstract

This invention provides a computing device and its operating method. The computing device includes multiple execution units, a global cache, and a data memory access unit. Each execution unit includes a first type of computing core, a second type of computing core, and a shared cache. The execution unit executes a thread group. The thread group includes multiple thread groups. The first thread group in the thread group controls the data memory access unit to transfer data between the global cache and the shared cache. Other thread groups in the thread group use the first type of computing core and the second type of computing core to perform calculations on the data located in the shared cache. The first operating time period of the first thread group partially overlaps with the second operating time periods of other thread groups to run in parallel. This invention's computing device and its operating method enable different types of computing cores in the same execution unit to execute in parallel, thereby making full use of computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a computing device in the field of artificial intelligence chips, and more particularly to a computing device and its operating method. Background Technology

[0002] The implementation of Artificial Intelligence (AI) applications requires enormous computing power from computing devices. This immense computing power originates from the numerous execution units (EUs, or execution cores) implemented in hardware within the computing device. Each execution unit can contain multiple computational cores (CUs), such as tensor cores, vector cores, and other computational cores. Other computational cores include integer (INT) computational cores or floating-point (FP) computational cores. By organizing these various types of computational cores through programming, execution units can support general-purpose computing, scientific computing, and neural network computing.

[0003] AI chips can execute groups of threads called warps using a Single Instruction, Multiple Thread (SIMT) approach. A warp is a group of threads. Many AI chip programming applications can leverage warps to achieve high performance. In warp-level programming, each warp handles the entire data transfer and computation from start to finish for different data sets. Therefore, in current warp-level programming and many computational task scenarios, tensor kernels and vector kernels can only be executed in a time-sharing manner (tensor kernels and vector kernels cannot be executed concurrently). Summary of the Invention

[0004] This invention relates to a computing device and its operating method, which enables different types of computing cores in the same execution unit to execute in parallel, thereby making full use of computing resources.

[0005] According to an embodiment of the present invention, a computing device includes a plurality of execution units, a global cache, and a data memory access unit. Each of the execution units includes a first type of computing core, a second type of computing core, and a shared cache. The data memory access unit is coupled to the global cache and the shared cache in the execution unit. The execution units execute thread groups, which include a plurality of thread groups, each thread group including a plurality of threads. A first thread group in the plurality of thread groups controls the data memory access unit to transfer data between the global cache and the shared cache. Other thread groups in the plurality of thread groups perform calculations on the data located in the shared cache using the first type of computing core and the second type of computing core. A first operating time period of the first thread group partially overlaps with the second operating time periods of the other thread groups to run in parallel.

[0006] According to an embodiment of the present invention, the computing device includes a plurality of execution units, a global cache, and a data memory access unit. Each of the execution units includes a first type of computing core, a second type of computing core, and a shared cache. The data memory access unit is coupled to the global cache and the shared cache of the execution unit. The method of operating the computing device includes: executing a thread group group through the execution units, wherein the thread group group includes a plurality of thread groups, and each thread group includes a plurality of threads. Executing the thread group group includes: executing a first thread group among the plurality of thread groups, the first thread group controlling the data memory access unit to transfer data between the global cache and the shared cache; and executing other thread groups among the plurality of thread groups, the other thread groups using the first type of computing core and the second type of computing core to perform calculations on the data located in the shared cache. A first operating time period of the first thread group and a second operating time period of the other thread groups partially overlap to run in parallel.

[0007] Based on the above, the computing device and its operation method described in this embodiment of the invention enable each execution unit to execute a thread group. The first thread group in this thread group is primarily used for data transmission, while other thread groups simultaneously or overlappingly perform calculations on the data using different types of computational cores (e.g., tensor cores and vector cores). In other words, this embodiment of the invention fully utilizes computing resources by enabling different types of computational cores within the same execution unit to execute in parallel. Attached Figure Description

[0008] Figure 1 This is a functional block diagram of a computing device according to an embodiment of the present invention.

[0009] Figure 2 This is a schematic diagram of the operation of each thread group in a computing device according to an embodiment of the present invention.

[0010] Figure 3 This is a flowchart of an operation method of a computing device according to an embodiment of the present invention.

[0011] Explanation of reference numerals in the attached figures

[0012] 100: Computing device

[0013] 105: Data Memory Access Unit

[0014] 107: Global Cache

[0015] 110-1~110-4: Execution Unit

[0016] 120-1~120-4: Type I computational kernels / tensor kernels

[0017] 122-1~122-4: Type II computational kernels / vector kernels

[0018] 130-1~130-4: Shared Cache

[0019] 210-1~210-2, 220-1~220-3, 230-1~230-3: Work tiles

[0020] S310~S330: Steps in the operation of the computing device

[0021] WG1: First Thread Group

[0022] WG2: Second Thread Group

[0023] WG3: The Third Thread Group

[0024] TMA: Data Transport Procedure

[0025] WGMMA: Matrix Multiply-Accumulate (MMA) Calculation Program for Thread Groups

[0026] EPI: Epilogue Calculation Procedure

[0027] TP1, TP2: Time periods Detailed Implementation

[0028] Reference will now be made in detail to exemplary embodiments of the invention, examples of which are illustrated in the accompanying drawings. Wherever possible, the same component symbols are used in the drawings and description to denote the same or similar parts.

[0029] The term "coupled (or connected)" as used throughout this specification (including the claims) may refer to any direct or indirect means of connection. For example, if the text describes a first device coupled (or connected) to a second device, it should be interpreted as the first device being directly connected to the second device, or the first device being indirectly connected to the second device through other devices or some means of connection. The terms "first," "second," etc., used throughout this specification (including the claims) are used to name components or distinguish different embodiments or scopes, and are not intended to limit the upper or lower limit of the number of components, nor to limit the order of components. Furthermore, wherever possible, components / components / steps using the same reference numerals in the drawings and embodiments represent the same or similar parts. Components / components / steps using the same reference numerals or the same terms in different embodiments may be referred to mutually in the relevant descriptions.

[0030] Figure 1 This is a functional block diagram of a computing device 100 according to an embodiment of the present invention. Figure 1 The computing device 100 includes multiple execution units (e.g., Figure 1 Execution units 110-1, 110-2, 110-3, and vector core 110-4; global cache 107; and data memory access unit 105. Execution units (e.g., Figure 1 Each of the execution units 110-1 includes a first type of computational core (e.g., tensor core 120-1), a second type of computational core (e.g., vector core 122-1), and a shared cache 130-1. The execution unit in this embodiment can have different types of computational cores. This embodiment uses tensor cores 120-1, 120-2, 120-3, and 120-4 as examples of the first type of computational core, and vector cores 122-1, 122-2, 122-3, and 122-4 as examples of the second type of computational core. In addition to the basic tensor and vector cores, it may also include integer computational cores, floating-point computational cores, etc. Users of this example can add other types of computational cores as needed.

[0031] Shared cache 130-1 is a cache shared by all computing cores in execution unit 110-1. Global cache 107 is a cache used by computing device 100 to transfer data with other external devices. Data memory access unit 105 is coupled to global cache 107 and execution unit (e.g., ...). Figure 1 Shared cache in execution units 110-1~110-4 (e.g., Figure 1Shared caches 130-1, 130-2, 130-3, and 130-4. Data memory access unit 105 is used to transfer data between these caches. Data memory access unit 105 is used to access data and is not used for instruction fetching.

[0032] The data memory access unit 105 in this embodiment can be implemented by a Tensor Memory Accelerator (TMA). The TMA can transfer large data blocks and multi-dimensional tensors from the global cache 107 to the shared caches 130-1 to 130-4, and transfer the calculated results back from the shared caches 130-1 to 130-4 to the global cache 107. The TMA can significantly reduce addressing overhead and improve efficiency, while supporting different tensor configurations, different cache access modes, and other functions.

[0033] Figure 1 The hardware structure of the computing device 100 is presented. The software programming of the computing device 100 is described below. The smallest unit in the software is described as a 'thread', but since each execution unit's computing core can handle multiple threads simultaneously, executing only a single thread would significantly reduce computational throughput. Therefore, multiple threads are integrated into a 'thread group' (also called a thread warp), and generally, the execution unit performs data transfer and computation based on one thread group after another. In this embodiment, 32 threads can be integrated into a thread group, and users of this embodiment can adjust the number of threads in a thread group according to their needs. However, if the execution unit executes only one thread group at a time, only one computing core of different types in the execution unit will be running, while the other computing cores will be idle. In other words, computing resources may not be fully utilized.

[0034] This invention enables each execution unit to execute one thread group at a time, and this thread group includes multiple thread groups. Each thread group includes multiple threads. The first thread group in the thread group is mainly used for data transfer. That is, the first thread group controls the data memory access unit 105 to transfer data between the global cache 107 and the shared caches 130-1 to 130-4. Other thread groups in the thread group control different types of computing cores (e.g., tensor cores and vector cores) to simultaneously or overlappingly perform calculations on the data in the shared caches 130-1 to 130-4, enabling different types of computing cores in the same execution unit to execute in parallel, thereby making full use of computing resources. In other words, the first operation time period of the first thread group partially overlaps with the second operation time periods of other thread groups to run in parallel.

[0035] Figure 2This is a schematic diagram illustrating the operation of various thread groups in a computing device according to an embodiment of the present invention. This embodiment uses... Figure 1 Execution unit 110-1 is used as an example for illustration. Figure 2 The execution unit 110-1 runs the aforementioned thread group. In this embodiment, the thread group may include N thread groups, where N is a positive integer and N is greater than or equal to 3. In this embodiment, N is set to 3, therefore these three thread groups are referred to as the first thread group WG1 and other thread groups (e.g., the second thread group WG2 and the third thread group WG3). Figure 2 The horizontal axis represents time. Figure 2 The working bricks 210-1~210-2, 220-1~220-3, and 230-1~230-3 in the first thread group WG1, the second thread group WG2, and the third thread group WG3 are presented as examples to illustrate the working procedure of these thread groups.

[0036] Figure 2 In the first thread group WG1, working blocks 210-1, 220-1, and 230-1 are all data transport programs (TMAs), that is, they control the data memory access unit 105 to transfer data between the global cache 107 and the shared caches 130-1 to 130-4. (Refer to...) Figure 1 and Figure 2 In this embodiment, each execution unit runs a thread group consisting of three thread groups (e.g., thread group WG1, thread group WG2, and thread group WG3). Each thread group contains M threads, where M is a positive integer and M is greater than or equal to 3. In this embodiment, M is set to 4, meaning that each thread group contains 4 threads. Therefore, the entire thread group can contain a total of 12 threads.

[0037] The first thread group WG1 may include four threads, which will be referred to here as thread one through thread four. (See reference...) Figure 1 The first thread is executed to move the input data from the global cache 107 to the shared cache 130-1. The second thread is executed to move the weighting data corresponding to the input data from the global cache 107 to the shared cache 130-1. The third thread is executed to move the result calculated by computation cores 120-1 and 122-1 from the shared cache 130-1 to the global cache 107. The fourth thread in this embodiment is an idle thread, that is, the fourth thread is reserved as a backup thread.

[0038] The second thread group WG2 and the third thread group WG3 each consist of four threads. Figure 2 In the second thread group WG2 and the third thread group WG3, one of them controls tensor kernel 120-1 to run the thread group matrix multiplication and accumulation (MMA) calculation program WGMMA on the data in shared cache 130-1, such as... Figure 2 Working bricks 210-2, 220-3, and 230-2 are shown. At the same time, one of the second thread group WG2 and the third thread group WG3 runs the epilogue calculation program EPI on the data in the shared cache 130-1 via control vector core 122-1, as shown... Figure 2 Working bricks 220-2 and 230-3 are shown. In conclusion, the EPI calculation program is run by vector kernel 122-1. The EPI calculation program can apply a bias, and then perform a Rectified Linear Unit (ReLU) vector transformation or apply / calculate the bias gradient on the matrix of input data. In other words, the first operating time period of the first thread group WG1 (e.g., ...) Figure 2 The second operating time period (e.g., time period TP1) and other thread groups (e.g., the second thread group WG2) Figure 2 The time periods TP2 will partially overlap to run in parallel.

[0039] exist Figure 2 During time period TP1, the first thread group WG1 will... Figure 1 The input data in global cache 107 is transferred to shared cache 130-1. After a delay period starting from time period TP1, the second thread group WG2 runs the thread group MMA calculation program WGMMA through tensor kernel 120-1 and stores the calculation results in shared cache 130-1.

[0040] exist Figure 2 During the time period corresponding to working brick 220-1, the first thread group WG1 continuously... Figure 1 Input data in global cache 107 is transferred to another part of shared cache 130-1. The second thread group WG2 runs the conclusion calculation program EPI through vector kernel 122-1, while the third thread group WG3 runs the thread group MMA calculation program WGMMA through tensor kernel 120-1 to compute the data in the other part of shared cache 130-1. This process repeats, with the second thread group WG2 and the third thread group WG3 simultaneously or overlappingly computing the data using different types of computational kernels (e.g., tensor kernels and vector kernels) (i.e., parallel execution), thus making fuller use of computational resources.

[0041] Figure 3 This is a flowchart of an operation method of a computing device according to an embodiment of the present invention. Figure 3 Operating method applicable Figure 1 Computing device 100. (See also...) Figure 1 and Figure 3 In step S310, the execution unit (e.g., Figure 1Execution unit 110-1) Execution thread group. The aforementioned thread group includes multiple thread groups, and each thread group includes multiple threads.

[0042] Step S310 includes steps S320 and S330. In step S320, the first thread group in the thread group is executed by the execution unit. The first thread group is executed by controlling the data memory access unit (e.g., ...). Figure 1 Data memory access unit 105) is used in global cache 107 and shared cache (e.g., Figure 1 Data is transferred between shared caches 130-1. In step S330, other thread groups in the thread group are executed by the execution unit. These other thread groups are processed through a first type of computing core (e.g., Figure 1 Tensor kernel 120-1) and second-type computational kernel (e.g., Figure 1 Vector kernel 122-1) to target the shared cache (e.g., Figure 1 The data in the shared cache (130-1) is used for calculation. The first operation time period of the first thread group partially overlaps with the second operation time period of other thread groups to run in parallel. In other words, steps S320 and S330 are performed simultaneously, and the order of the steps is not necessarily fixed.

[0043] In summary, the computing device and its operation method described in the embodiments of the present invention enable each execution unit to execute a group of threads. The first thread group in this group is primarily used for data transmission, while other thread groups simultaneously or overlappingly perform calculations on the data using different types of computational cores (e.g., tensor cores and vector cores). In other words, the embodiments of the present invention fully utilize computing resources by enabling different types of computational cores within the same execution unit to execute in parallel.

[0044] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A computing device, characterized in that, include: Multiple execution units, wherein each execution unit includes a first type of computing core, a second type of computing core, and a shared cache; Global caching; as well as The data memory access unit is coupled to the global cache and the shared cache in the execution unit. The execution unit executes a thread group group, which includes multiple thread groups, and each thread group includes multiple threads. The first thread group among the plurality of thread groups controls the data memory access unit to transfer data between the global cache and the shared cache. Other thread groups in the plurality of thread groups perform calculations on the data located in the shared cache using the first type of computing core and the second type of computing core. The first operating time period of the first thread group used for data transmission partially overlaps with the second operating time period of the other thread groups used for computation, so that they can run in parallel. The first type of computational kernel is a tensor kernel, and the second type of computational kernel is a vector kernel; The other thread groups include a second thread group and a third thread group, wherein one of the second thread group and the third thread group controls the tensor kernel to perform matrix multiplication and accumulation calculations on the data in the shared cache, and at the same time, the other of the second thread group and the third thread group controls the vector kernel to perform conclusion calculations on the data in the shared cache, wherein the conclusion calculations apply a bias and then perform linear rectified unit function vector transformation or bias gradient calculations on the matrix of the input data; In this process, while the second thread group is performing the matrix multiplication and accumulation calculation through the tensor kernel, the third thread group is performing the conclusion calculation through the vector kernel. Furthermore, while the second thread group is performing the conclusion calculation through the vector kernel, the third thread group is performing the matrix multiplication and accumulation calculation through the tensor kernel. The second thread group performs different calculations within the time periods corresponding to two adjacent working blocks in the first thread group, and the third thread group performs different calculations within the time periods corresponding to two adjacent working blocks in the first thread group. In this context, the working brick in the first thread group is a data transport program.

2. The computing device according to claim 1, characterized in that, The first thread group includes a first thread, a second thread, and a third thread. The first thread is executed to move the input data from the global cache to the shared cache. The second thread is executed to move the weight data corresponding to the input data from the global cache to the shared cache. The third thread is executed to move the calculated result from the shared cache to the global cache.

3. The computing device according to claim 2, characterized in that, The first thread group also includes a fourth thread, wherein the fourth thread is an idle thread.

4. The computing device according to claim 1, characterized in that, The thread group comprises N thread groups, where N is a positive integer and N is greater than or equal to 3. Each thread group consists of M threads, where M is a positive integer and M is greater than or equal to 3.

5. The computing device according to claim 1, characterized in that, The data memory access unit is a tensor memory accelerator.

6. A method for operating a computing device, characterized in that, The computing device includes multiple execution units, a global cache, and a data memory access unit. Each execution unit includes a first type of computing core, a second type of computing core, and a shared cache. The data memory access unit is coupled to the global cache and the shared cache of the execution unit. The operation method includes: The execution unit executes thread groups, wherein the thread groups include multiple thread groups, and each thread group includes multiple threads. Executing the thread group includes: The first thread group of the plurality of thread groups is executed, and the first thread group controls the data memory access unit to transfer data between the global cache and the shared cache; and Other thread groups among the plurality of thread groups are executed, and these other thread groups perform calculations on the data located in the shared cache using the first type of computing core and the second type of computing core. The first operating time period of the first thread group used for data transmission partially overlaps with the second operating time period of the other thread groups used for computation, so that they can run in parallel. The first type of computational kernel is a tensor kernel, and the second type of computational kernel is a vector kernel; The other thread groups include a second thread group and a third thread group, wherein one of the second thread group and the third thread group controls the tensor kernel to perform matrix multiplication and accumulation calculations on the data in the shared cache, and at the same time, the other of the second thread group and the third thread group controls the vector kernel to perform conclusion calculations on the data in the shared cache, wherein the conclusion calculations apply a bias and then perform linear rectified unit function vector transformation or bias gradient calculations on the matrix of the input data; In this process, while the second thread group is performing the matrix multiplication and accumulation calculation through the tensor kernel, the third thread group is performing the conclusion calculation through the vector kernel. Furthermore, while the second thread group is performing the conclusion calculation through the vector kernel, the third thread group is performing the matrix multiplication and accumulation calculation through the tensor kernel. The second thread group performs different calculations within the time periods corresponding to two adjacent working blocks in the first thread group, and the third thread group performs different calculations within the time periods corresponding to two adjacent working blocks in the first thread group. In this context, the working brick in the first thread group is a data transport program.

7. The operating method according to claim 6, characterized in that, The first thread group includes a first thread, a second thread, and a third thread. The first thread is executed to move the input data from the global cache to the shared cache. The second thread is executed to move the weight data corresponding to the input data from the global cache to the shared cache. The third thread is executed to move the calculated result from the shared cache to the global cache.

Citation Information

Patent Citations

  • Arithmetic device, method for operating arithmetic device, and machine-readable storage medium

    CN117992235A

  • Data processor, data processing method, electronic equipment and storage medium

    CN118035618A