Artificial intelligence chip and operation method thereof
By introducing dedicated tensor and general-purpose thread bundle scheduling and instruction issuance units into the artificial intelligence chip, asynchronous execution of tensor cores and general-purpose computing cores is achieved, solving the problem of low efficiency of synchronous execution and improving the computing efficiency of the chip.
Patent Information
- Application Number
- CN202511157205.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-08-19
AI Technical Summary
In the existing technology, general-purpose computing cores and tensor cores share the same thread bundle scheduling and instruction emission unit, resulting in asynchronous instruction execution and low execution efficiency. In particular, when the general-purpose computing core does not exit in time, the tensor core cannot start the next thread bundle, resulting in a waste of computing power.
Multiple tensor thread warp scheduling and instruction emission units and multiple general thread warp scheduling and instruction emission units are introduced. The thread blocks are divided into tensor computing and non-tensor computing thread warps through the thread block splitting unit, and dispatched to dedicated tensor cores and general computing cores respectively to achieve asynchronous execution.
Through asynchronous execution, the execution efficiency of artificial intelligence chips is improved, the waste of computing power caused by synchronous dependence is avoided, and the overall computing power is improved.
Smart Images

Figure CN120655494A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of integrated circuit technology, and in particular to an artificial intelligence (AI) chip and an operating method thereof. Background Art
[0002] A warp is the basic execution unit in a graphics processing unit (GPU), general-purpose graphics processing unit (GPGPU) chip, or artificial intelligence chip. Generally speaking, each warp contains 32 threads. When the number of threads in a user-defined thread group or thread block is not an integer multiple of 32, the thread block is assigned to N+1 warps, where N is an integer. A warp is a collection of threads at the hardware level, while a thread block is a collection of threads at the logical level. The threads in each warp execute using a single instruction multiple thread (SIMT) mechanism, meaning that threads in the same warp all execute the same instruction.
[0003] General-purpose compute cores and tensor cores share these thread warps. Because a single thread warp can only execute the same instruction stream (assembler), the instruction issue unit issues instructions to the general-purpose compute cores and tensor cores sequentially according to the instruction stream order. In other words, the instruction issue unit cannot send instructions to both general-purpose compute cores and tensor cores simultaneously; it can only issue general-purpose compute instructions, tensor compute instructions, and other instructions serially. General-purpose compute cores and tensor cores are physically independent. Under existing technology, general-purpose compute cores and tensor cores share the same thread warp scheduling and instruction issue unit, resulting in incomplete asynchronous execution of their respective instructions. General-purpose compute cores and tensor cores may block each other, reducing execution efficiency. For example, if the thread warp being computed by the general-purpose compute core cannot exit in a timely manner, the tensor core cannot start the next thread warp, resulting in wasted computing power, known as the long-tail effect. Improving execution efficiency is one of the many technical issues in the field of AI chips. Summary of the Invention
[0004] The present invention provides an artificial intelligence chip and an operating method thereof to improve execution efficiency.
[0005] In an embodiment according to the present invention, an artificial intelligence chip includes multiple tensor cores, multiple tensor warp scheduling and instruction issuing units (the tensor warp scheduling and instruction issuing units may also be referred to as tensor issuing units), multiple general-purpose computing cores, multiple general-purpose warp scheduling and instruction issuing units (the general-purpose warp scheduling and instruction issuing units may also be referred to as general-purpose issuing units), and a thread block splitting unit. A first tensor warp scheduling and instruction issuing unit among the multiple tensor warp scheduling and instruction issuing units is coupled to a first tensor core among the multiple tensor cores. A first general-purpose warp scheduling and instruction issuing unit among the multiple general-purpose computing cores is coupled to the first tensor core and the first general-purpose computing core among the multiple general-purpose computing cores. A thread block splitting unit is coupled to the multiple tensor warp scheduling and instruction issuing units and the multiple general-purpose warp scheduling and instruction issuing units. The thread block splitting unit splits a current thread block into multiple warps. In response to the thread block slicing unit operating in a first operating mode, the thread block slicing unit dispatches each of the plurality of warps to one of the plurality of general-purpose warp scheduling and instruction issuing units, and the first general-purpose warp scheduling and instruction issuing unit issues each instruction of the current warp to one of the first tensor core and the first general-purpose computing core. In response to the thread block slicing unit operating in a second operating mode, the thread block slicing unit dispatches each of at least one tensor warp involved in tensor computation among the plurality of warps to one of the plurality of tensor warp scheduling and instruction issuing units, the first tensor warp scheduling and instruction issuing unit issues the tensor computation instruction of the current tensor warp to the first tensor core, the thread block slicing unit dispatches each of at least one non-tensor warp not involved in tensor computation among the plurality of warps to one of the plurality of general-purpose warp scheduling and instruction issuing units, and the first general-purpose warp scheduling and instruction issuing unit issues the non-tensor computation instruction of the current non-tensor warp to the first general-purpose computing core.
[0006] In an embodiment according to the present invention, the operation method includes: dividing a current thread block into a plurality of thread warps by a thread block slicing unit of an artificial intelligence chip; in response to the thread block slicing unit operating in a first operation mode, dispatching each of the plurality of thread warps to one of a plurality of general thread warp scheduling and instruction issuing units by the thread block slicing unit, and issuing each instruction of the current thread warp to one of a first tensor core and a first general computing core by the first general thread warp scheduling and instruction issuing unit; and in response to the thread block slicing unit operating in a second operation mode, dispatching the plurality of thread warps to one of a plurality of general thread warp scheduling and instruction issuing units by the thread block slicing unit. Each of at least one tensor thread warp involved in tensor computation in the thread warp is dispatched to one of the multiple tensor thread warp scheduling and instruction issuing units, the first tensor thread warp scheduling and instruction issuing unit issues the tensor computation instruction of the current tensor thread warp to the first tensor core, the thread block splitting unit dispatches each of at least one non-tensor thread warp not involved in tensor computation in the multiple thread warps to one of the multiple general-purpose thread warp scheduling and instruction issuing units, and the first general-purpose thread warp scheduling and instruction issuing unit issues the non-tensor computation instruction of the current non-tensor thread warp to the first general-purpose computing core.
[0007] Based on the above, the thread block splitting unit can dispatch tensor thread bundles involving tensor calculations to the tensor thread bundle scheduling and instruction emission unit, and dispatch non-tensor thread bundles not involving tensor calculations to the general thread bundle scheduling and instruction emission unit. Based on this, when the tensor thread bundle scheduling and instruction emission unit emits the tensor calculation instructions of the tensor thread bundle to the tensor core, the general thread bundle scheduling and instruction emission unit can simultaneously emit the non-tensor calculation instructions of the non-tensor thread bundle to the general computing core. In other words, the tensor core and the general computing core can perform their respective computing operations at the same time. Therefore, in the second operating mode, the execution efficiency of the artificial intelligence chip can be effectively improved.
[0008] In order to make the above features and advantages of the present invention more clearly understood, embodiments are given below with reference to the accompanying drawings for detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 is a schematic diagram of a circuit block of an artificial intelligence chip according to one embodiment; Figure 2 1 is a schematic diagram of a circuit module of an artificial intelligence chip according to an embodiment of the present invention; Figure 3 This is a flowchart of an operating method of an artificial intelligence chip according to one embodiment of the present invention; Figure 4 FIG. 1 is a schematic diagram of a circuit module of a tensor warp scheduling and instruction issuing unit according to an embodiment of the present invention.
[0010] Explanation of Figure Numbers 100: No. 1 artificial intelligence chip; 200: 2nd artificial intelligence chip; 110: 1st thread block splitting unit; 210: second thread block splitting unit; 120_1: Thread warp 1 scheduling and instruction issuing unit; 120_N: Nth warp scheduling and instruction issuing unit; 130: 1st warp synchronization unit; 240: 2nd warp synchronization unit; 140_1: 1_1th tensor core; 140_N: 1_Nth tensor core; 250: Tensor Core; 250_1: 2_1st tensor core; 250_N: 2_Nth tensor core; 150_1: 1st general computing core; 150_N: 1_Nth general computing core; 260_1: 2_1 general computing core; 260_N: 2_Nth general computing core; 160: 1st shared memory area; 270: 2nd shared memory area; 211: Independent scheduling switch register; 220: tensor thread warp scheduling and instruction issuance unit; 220_1: 1st tensor warp scheduling and instruction issuing unit; 220_N: Nth tensor warp scheduling and instruction issuing unit; 221: buffer; 222: thread warp scheduling unit; 223: instruction decoding unit; 224: operand read unit; 225: scalar instruction execution unit; 226: scalar register group; 227: Tensor Core instruction forwarding unit; 230_1: 1st general thread warp scheduling and instruction issuing unit; 230_N: Nth general warp scheduling and instruction issuing unit. DETAILED DESCRIPTION
[0011] Reference will now be made in detail to exemplary embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Whenever possible, the same reference numerals are used in the drawings and the description to refer to the same or like parts.
[0012] The term "coupled (or connected)" as used throughout this specification (including the claims) may refer to any direct or indirect means of connection. For example, if a first device is described as being coupled (or connected) to a second device, this should be interpreted to mean that the first device may be directly connected to the second device, or the first device may be indirectly connected to the second device via another device or some other means of connection. Terms such as "first" and "second" throughout this specification (including the claims) are used to name components or to distinguish between different embodiments or scopes, and are not intended to limit the upper or lower limits on the number of components or the order of components. Furthermore, wherever possible, components, members, and steps using the same reference numbers in the drawings and embodiments represent identical or similar parts. Components, members, and steps using the same reference numbers or the same terminology in different embodiments may refer to the relevant descriptions. It should be understood that features of the following embodiments may be combined. For example, features of the second embodiment may be implemented in combination with features of the first embodiment. Persons with ordinary skill in the art can select appropriate feature combinations based on actual design requirements.
[0013] Computing devices such as artificial intelligence (AI) chips can provide enormous computing power. This immense computing power stems from their numerous internal hardware cores. An AI chip typically contains multiple programmable processors, such as a stream processor cluster (SPC). Each programmable processor typically includes multiple compute units (CUs, or cores), and each CU typically includes multiple execution units (EUs, or cores), such as tensor cores (Tcores) and general-purpose cores. General-purpose cores typically include at least one of integer (INT) cores, floating-point (FP) cores, and vector cores (Vcores). By programming and organizing various types of compute units, AI chips can support general-purpose computing, scientific computing, and neural network computing.
[0014] Figure 1 2 is a schematic diagram of a circuit module of an artificial intelligence chip according to one embodiment. Figure 1 The first artificial intelligence chip 100 includes a first thread block splitting unit 110, a plurality of thread warp scheduling and instruction issuing units (e.g. Figure 1, . . . , the Nth warp scheduling and instruction issuing unit 120_N), the first warp synchronization unit 130, a plurality of tensor cores (eg Figure 1 1_1 th tensor core 140_1, ..., 1_N th tensor core 140_N), multiple general computing cores (e.g. Figure 1 1_1 general-purpose computing cores 150_1, ..., 1_N general-purpose computing cores 150_N) and a first shared memory area 160. Based on actual design and application, 1_1 general-purpose computing cores 150_1 through 1_N general-purpose computing cores 150_N include at least one of an integer core, a floating-point core, and a vector core. The first warp scheduling and instruction-issuing unit 120_1 is coupled to the first tensor core 140_1 and the first general-purpose computing core 150_1. Similarly, the Nth warp scheduling and instruction-issuing unit 120_N is coupled to the first tensor core 140_N and the first general-purpose computing core 150_N. The number N of warp scheduling and instruction-issuing units 120_1 through 120_N can be determined based on actual design and application.
[0015] The first thread block splitting unit 110 is coupled to the first through Nth warp scheduling and instruction-issuing units 120_1 and 120_N. The thread block splitting unit (first thread block splitting unit 110) splits the current thread block into multiple warps. Based on user (software)-specified rules, the first thread block splitting unit 110 dispatches all warps to the first through Nth warp scheduling and instruction-issuing units 120_1 and 120_N. For example, the first thread block splitting unit 110 assigns the warps of a thread block to the first through Nth warp scheduling and instruction-issuing units 120_1 and 120_N in a round-robin or fixed order. The first thread block splitting unit 110 dispatches each warp to one of the first through Nth warp scheduling and instruction-issuing units 120_1 and 120_N. The 1st warp scheduling and instruction-issuing unit 120_1 issues each instruction of the local warp to one of the 1_1th tensor core 140_1 and the 1_1th general-purpose computing core 150_1. For example, the 1st warp scheduling and instruction-issuing unit 120_1 issues tensor computation instructions involving tensor computations to the 1_1th tensor core 140_1 and non-tensor computation instructions not involving tensor computations to the 1_1th general-purpose computing core 150_1. Similarly, the Nth warp scheduling and instruction-issuing unit 120_N issues tensor computation instructions of the local warp to the 1_Nth tensor core 140_N and non-tensor computation instructions of the local warp to the 1_Nth general-purpose computing core 150_N.
[0016] The first warp synchronization unit 130 is coupled to the first through Nth warp scheduling and instruction-issuing units 120_1 and 120_N to receive synchronization instructions programmed by users in the warps. Based on the synchronization instructions, the first warp synchronization unit 130 selectively controls instruction issuance by one or more of the first through Nth warp scheduling and instruction-issuing units 120_1 and 120_N. A first shared memory area 160 is coupled to the 1_1st through 1_Nth tensor cores 140_1 and 1_1st through 1_Nth general-purpose computing cores 150_1 and 1_1st through 1_Nth general-purpose computing cores 150_N. The 1_1st through 1_Nth tensor cores 140_1 and 1_1st through 1_Nth general-purpose computing cores 150_1 and 1_1st through 1_Nth general-purpose computing cores 150_N can exchange data in the first shared memory area 160.
[0017] The first thread block splitting unit 110 assigns multiple warps within a thread block to the first through Nth warp scheduling and instruction-issuing units 120_1 through 120_N in a round-robin or fixed order. These warps are shared by the first-first general-purpose computing cores 150_1 through 1_Nth general-purpose computing cores 150_1 through 1_Nth general-purpose computing cores 150_N and the first-first tensor cores 140_1 through 1_Nth tensor cores 140_N. Because a warp can only execute the same instruction stream (assembler), the first through Nth warp scheduling and instruction-issuing units 120_1 through 120_N issue instructions to the first-first general-purpose computing cores 150_1 through 1_Nth general-purpose computing cores 150_N and the first-first tensor cores 140_1 through 1_Nth tensor cores 140_N, following the order of the instruction streams. In other words, the instruction issuing device (each of the 1st thread warp scheduling and instruction issuing unit 120_1 to the Nth thread warp scheduling and instruction issuing unit 120_N) cannot send instructions to the general computing core and the tensor core at the same time (for example, the 1st thread warp scheduling and instruction issuing unit 120_1 cannot send instructions to the 1_1th general computing core 150_1 and the 1_1th tensor core 140_1 at the same time), and can only serially issue non-tensor computing instructions (general computing instructions), tensor computing instructions, and other instructions.
[0018] The 1_1th general-purpose computing core 150_1 and the 1_1th tensor core 140_1 are physically independent and share the 1_1th warp scheduling and instruction issuance unit 120_1. If the warp being computed by the 1_1th general-purpose computing core 150_1 fails to exit in a timely manner, the 1_1th tensor core 140_1 cannot start the next warp, resulting in wasted computing power and reduced execution efficiency. The same applies to the N_th warp scheduling and instruction issuance unit 120_N, the 1_Nth tensor core 140_N, and the 1_Nth general-purpose computing core 150_N.
[0019] The following embodiments decouple tensor cores from general-purpose computing cores by introducing dedicated warp scheduling and instruction issuance units for tensor cores. Logically, the warps of tensor cores and general-purpose computing cores still belong to the same thread block and execute the same program, but they jump to their respective program segments based on the warp sequence number.
[0020] Figure 2 FIG. 1 is a schematic diagram of a circuit module of an artificial intelligence chip according to an embodiment of the present invention. Figure 2 In the embodiment of the present invention, the second artificial intelligence chip 200 includes a second thread block splitting unit 210, a plurality of tensor warp scheduling and instruction issuing units (which may be referred to as tensor issuing units, for example Figure 2 , . . . , the Nth tensor warp scheduling and instruction issuing unit 220_N), a plurality of general-purpose warp scheduling and instruction issuing units (which may be referred to as general-purpose issuing units, for example Figure 2 , . . . , the Nth general warp scheduling and instruction issuing unit 230_1 , . . . , the Nth general warp scheduling and instruction issuing unit 230_N), the second warp synchronization unit 240 , a plurality of tensor cores (eg Figure 2 2_1 th tensor core 250_1, ..., 2_N th tensor core 250_N), multiple general computing cores (e.g. Figure 2 , 2_1 th general computing core 260_1, ..., 2_N th general computing core 260_N) and a second shared memory area 270. The number N of the 1 th tensor warp scheduling and instruction issuing unit 220_1 to the N th tensor warp scheduling and instruction issuing unit 220_N may be determined according to actual design and application.
[0021] Among them, any one of the 1st tensor thread warp scheduling and instruction issuance unit 220_1 to the Nth tensor thread warp scheduling and instruction issuance unit 220_N can be equivalent to the first tensor thread warp scheduling and instruction issuance unit, and any other one of the 1st tensor thread warp scheduling and instruction issuance unit 220_1 to the Nth tensor thread warp scheduling and instruction issuance unit 220_N can be equivalent to the second tensor thread warp scheduling and instruction issuance unit, for example, the 1st tensor thread warp scheduling and instruction issuance unit 220_1 can be equivalent to the first tensor thread warp scheduling and instruction issuance unit, and the Nth tensor thread warp scheduling and instruction issuance unit 220_N can be equivalent to the second tensor thread warp scheduling and instruction issuance unit.
[0022] Among them, any one of the 1st general thread warp scheduling and instruction issuing unit 230_1 to the Nth general thread warp scheduling and instruction issuing unit 230_N can be equivalent to the first general thread warp scheduling and instruction issuing unit, and any other one of the 1st general thread warp scheduling and instruction issuing unit 230_1 to the Nth general thread warp scheduling and instruction issuing unit 230_N can be equivalent to the second general thread warp scheduling and instruction issuing unit. For example, the 1st general thread warp scheduling and instruction issuing unit 230_1 can be equivalent to the first general thread warp scheduling and instruction issuing unit, and the Nth general thread warp scheduling and instruction issuing unit 230_N can be equivalent to the second general thread warp scheduling and instruction issuing unit.
[0023] Among them, any one of the 2_1th tensor core 250_1 to the 2_Nth tensor core 250_N can be equivalent to the first tensor core, and any other one of the 2_1th tensor core 250_1 to the 2_Nth tensor core 250_N can be equivalent to the second tensor core. For example, the 2_1th tensor core 250_1 can be equivalent to the first tensor core, and the 2_Nth tensor core 250_N can be equivalent to the second tensor core.
[0024] Among them, any one of the 2_1th general computing core 260_1 to the 2_Nth general computing core 260_N can be equivalent to the first general computing core, and any other one of the 2_1th general computing core 260_1 to the 2_Nth general computing core 260_N can be equivalent to the second general computing core. For example, the 2_1th general computing core 260_1 can be equivalent to the first general computing core, and the 2_Nth general computing core 260_N can be equivalent to the second general computing core.
[0025] According to different designs, in some embodiments, at least one of the above-mentioned second thread block splitting unit 210, the first tensor thread warp scheduling and instruction issuing unit 220_1 to the Nth tensor thread warp scheduling and instruction issuing unit 220_N, the first general thread warp scheduling and instruction issuing unit 230_1 to the Nth general thread warp scheduling and instruction issuing unit 230_N, the second warp synchronization unit 240, the 2_1st tensor core 250_1 to the 2_Nth tensor core 250_N, and the 2_1st general computing core 260_1 to the 2_Nth general computing core 260_N can be implemented as a hardware circuit. In other embodiments, at least one of the second thread block splitting unit 210, the first tensor warp scheduling and instruction issuing unit 220_1 to the Nth tensor warp scheduling and instruction issuing unit 220_N, the first general-purpose warp scheduling and instruction issuing unit 230_1 to the Nth general-purpose warp scheduling and instruction issuing unit 230_N, the second warp synchronization unit 240, the 2_1st tensor core 250_1 to the 2_Nth tensor core 250_N, and the 2_1st general-purpose computing core 260_1 to the 2_Nth general-purpose computing core 260_N may be implemented in a combination of hardware, firmware, and software (i.e., a program).
[0026] In hardware terms, at least one of the second thread block splitting unit 210, the first through Nth tensor warp scheduling and instruction issuing units 220_1 through 220_N, the first through Nth general-purpose warp scheduling and instruction issuing units 230_1 through 230_N, the second warp synchronization unit 240, the 2_1st through 2_Nth tensor cores 250_1 through 250_N, and the 2_1st through 2_Nth general-purpose computing cores 260_1 through 260_N can be implemented as a logic circuit on an integrated circuit. For example, at least one of the functions related to the second thread block splitting unit 210, the first tensor warp scheduling and instruction issuing unit 220_1 to the Nth tensor warp scheduling and instruction issuing unit 220_N, the first general-purpose warp scheduling and instruction issuing unit 230_1 to the Nth general-purpose warp scheduling and instruction issuing unit 230_N, the second warp synchronization unit 240, the 2_1st tensor core 250_1 to the 2_Nth tensor core 250_N, and the 2_1st general-purpose computing core 260_1 to the 2_Nth general-purpose computing core 260_N can be implemented in one or more hardware controllers, microcontrollers, hardware processors, microprocessors, application-specific integrated circuits (ASICs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), central processing units (CPUs), or other similar processors. Various logic blocks, modules, and circuits in a CPU or other processing unit. The functions associated with at least one of the second thread block splitting unit 210, the first through Nth tensor warp scheduling and instruction issuing units 220_1 through 220_N, the first through Nth general-purpose warp scheduling and instruction issuing units 230_1 through 230_N, the second warp synchronization unit 240, the 2_1st through 2_Nth tensor cores 250_1 through 250_N, and the 2_1st through 2_Nth general-purpose computing cores 260_1 through 260_N can be implemented as hardware circuits using hardware description languages (such as Verilog HDL or VHDL) or other suitable programming languages, such as various logic blocks, modules, and circuits in an integrated circuit.
[0027] In software or firmware form, functions related to at least one of the second thread block splitting unit 210, the first tensor warp scheduling and instruction issuing unit 220_1 to the Nth tensor warp scheduling and instruction issuing unit 220_N, the first general-purpose warp scheduling and instruction issuing unit 230_1 to the Nth general-purpose warp scheduling and instruction issuing unit 230_N, the second warp synchronization unit 240, the 2_1st tensor cores 250_1 to the 2_Nth tensor cores 250_N, and the 2_1st general-purpose computing cores 260_1 to the 2_Nth general-purpose computing cores 260_N may be implemented as programming codes. For example, at least one of the second thread block splitting unit 210 , the first through Nth tensor warp scheduling and instruction issuing units 220_1 through 220_N, the first through Nth general-purpose warp scheduling and instruction issuing units 230_1 through 230_N, the second warp synchronization unit 240 , the 2_1st through 2_Nth tensor cores 250_1 through 250_N, and the 2_1st through 2_Nth general-purpose computing cores 260_1 through 260_N may be implemented using general programming languages (e.g., C, C++, or assembly) or other suitable programming languages. The programming code may be recorded and stored in a non-transitory machine-readable storage medium. In some embodiments, the non-transitory machine-readable storage medium may include, for example, a semiconductor memory and / or a storage device. An electronic device (e.g., a CPU, a hardware controller, a microcontroller, a hardware processor, or a microprocessor) can read and execute programming code from a non-transitory machine-readable storage medium to implement the relevant functions of at least one of the second thread block splitting unit 210, the first tensor warp scheduling and instruction issuing unit 220_1 to the Nth tensor warp scheduling and instruction issuing unit 220_N, the first general-purpose warp scheduling and instruction issuing unit 230_1 to the Nth general-purpose warp scheduling and instruction issuing unit 230_N, the second warp synchronization unit 240, the 2_1st tensor core 250_1 to the 2_Nth tensor core 250_N, and the 2_1st general-purpose computing core 260_1 to the 2_Nth general-purpose computing core 260_N.
[0028] The second thread block splitting unit 210 includes an independent scheduling switch register 211. The state of the independent scheduling switch register 211 depends on the kernel descriptor. The kernel descriptor is generated by the driver. Generally speaking, the instruction binary file address, user configuration information, required hardware resources, and other information of the current executable program are packaged into a fixed-format structure as the unique identification information of the current executable program. This is the kernel descriptor. In response to the independent scheduling switch register 211 indicating that the switch is closed, the second thread block splitting unit 210 operates in the first operating mode. In response to the independent scheduling switch register 211 indicating that the switch is open, the second thread block splitting unit 210 operates in the second operating mode.
[0029] Figure 3 This is a flow chart of an operation method of an artificial intelligence chip according to an embodiment of the present invention. Figure 2 and Figure 3 In step S310, the second thread block splitting unit 210 splits the current thread block into a plurality of thread warps. The second thread block splitting unit 210 is coupled to the first tensor warp scheduling and instruction issuing unit 220_1 to the Nth tensor warp scheduling and instruction issuing unit 220_N and the first general warp scheduling and instruction issuing unit 230_1 to the Nth general warp scheduling and instruction issuing unit 230_N. In response to the second thread block splitting unit 210 operating in the first operation mode, the second thread block splitting unit 210 dispatches each thread warp to one of the first general warp scheduling and instruction issuing unit 230_1 to the Nth general warp scheduling and instruction issuing unit 230_N (step S320). The second thread block splitting unit 210 operating in the first operation mode can refer to Figure 1 The related description of the first thread block partitioning unit 110 is similar to that of the first thread block partitioning unit 110 and is not repeated here.
[0030] In response to the second thread block slicing unit 210 operating in the second operation mode, the second thread block slicing unit 210 dispatches each tensor warp involved in tensor computation among the multiple warps to one of the first to Nth tensor warp scheduling and instruction-issuing units 220_1 to 220_N, and dispatches each non-tensor warp not involved in tensor computation among the multiple warps to one of the first to Nth general warp scheduling and instruction-issuing units 230_1 to 230_N (step S340). In some application examples, a program descriptor may define and indicate tensor warps involved in tensor computation among multiple warps, and the second thread block splitting unit 210 dispatches each tensor warp to one of the first tensor warp scheduling and instruction issuing unit 220_1 to the Nth tensor warp scheduling and instruction issuing unit 220_N based on the program descriptor, and dispatches each non-tensor warp to one of the first general warp scheduling and instruction issuing unit 230_1 to the Nth general warp scheduling and instruction issuing unit 230_N.
[0031] The second warp synchronization unit 240 is coupled to the first tensor warp scheduling and instruction issuing unit 220_1 to the Nth tensor warp scheduling and instruction issuing unit 220_N and the first general warp scheduling and instruction issuing unit 230_1 to the Nth general warp scheduling and instruction issuing unit 230_N to receive the synchronization instruction. The second warp synchronization unit 240 selectively controls the instruction issuance of one or more of the first tensor warp scheduling and instruction issuing unit 220_1 to the Nth tensor warp scheduling and instruction issuing unit 220_N and the first general warp scheduling and instruction issuing unit 230_1 to the Nth general warp scheduling and instruction issuing unit 230_N based on the synchronization instruction. The second warp synchronization unit 240 can refer to Figure 1 The related description of the first warp synchronization unit 130 is similar and will not be repeated here.
[0032] The 1st tensor warp scheduling and instruction-issuing unit 220_1 is coupled to the 2_1st tensor core 250_1, and the 1st general-purpose warp scheduling and instruction-issuing unit 230_1 is coupled to the 2_1st tensor core 250_1 and the 2_1st general-purpose computing core 260_1. Similarly, the tensor warp scheduling and instruction-issuing unit 220_N is coupled to the 2_Nth tensor core 250_N, and the Nth general-purpose warp scheduling and instruction-issuing unit 230_N is coupled to the 2_Nth tensor core 250_N and the 2_Nth general-purpose computing core 260_N. Based on actual design and application, the 2_1st through 2_Nth general-purpose computing cores 260_1 through 260_N may include at least one of an integer core, a floating-point core, or a vector core. The 2_1st tensor core 250_1 to the 2_Nth tensor core 250_N and the 2_1st general computing core 260_1 to the 2_Nth general computing core 260_N can refer to Figure 1 The related descriptions of the 1_1th tensor core 140_1 to the 1_Nth tensor core 140_N and the 1_1th general computing core 150_1 to the 1_Nth general computing core 150_N are similar and will not be repeated here.
[0033] The first tensor warp scheduling and instruction-issuing unit 220_1, the first general-purpose warp scheduling and instruction-issuing unit 230_1, the second-first tensor core 250_1, and the second-first general-purpose computing core 260_1 are used as examples for illustration. The remaining tensor warp scheduling and instruction-issuing units (e.g., the Nth tensor warp scheduling and instruction-issuing unit 220_N), the remaining general-purpose warp scheduling and instruction-issuing units (e.g., the Nth general-purpose warp scheduling and instruction-issuing unit 230_N), the remaining tensor cores (e.g., the second-Nth tensor core 250_N), and the remaining general-purpose computing cores (e.g., the second-Nth general-purpose computing core 260_N) can refer to the description of the first tensor warp scheduling and instruction-issuing unit 220_1, the first general-purpose warp scheduling and instruction-issuing unit 230_1, the second-first tensor core 250_1, and the second-first general-purpose computing core 260_1, and the description thereof will be omitted for brevity. In response to the second thread block splitting unit 210 operating in the first operating mode, the second thread block splitting unit 210 dispatches the warp to the first general warp scheduling and instruction issuing unit 230_1 (step S320). Similarly, the second thread block splitting unit 210 dispatches the warp to the Nth general warp scheduling and instruction issuing unit 230_N (step S320). The first general warp scheduling and instruction issuing unit 230_1 issues each instruction of the current warp to one of the 2_1st tensor core 250_1 and the 2_1st general computing core 260_1 (step S330). Similarly, the Nth general warp scheduling and instruction issuing unit 230_N issues each instruction of the current warp to one of the 2_Nth tensor core 250_N and the 2_Nth general computing core 260_N (step S330). In the first operation mode, the first tensor thread warp scheduling and instruction issuing unit 220_1 to the Nth tensor thread warp scheduling and instruction issuing unit 220_N are idle. The second artificial intelligence chip 200 operating in the first operation mode can refer to Figure 1 The relevant description of the first artificial intelligence chip 100 is shown.
[0034] In response to the second thread block slicing unit 210 operating in the second operating mode, the second thread block slicing unit 210 dispatches tensor warps involved in tensor computations from the multiple warps to the first tensor warp scheduling and instruction issuing unit 220_1. The first tensor warp scheduling and instruction issuing unit 220_1 then issues the tensor computation instructions for the current tensor warp to the 2_1st tensor core 250_1 (step S340). In response to the second thread block slicing unit 210 operating in the second operating mode, the second thread block slicing unit 210 dispatches non-tensor warps not involved in tensor computations from the multiple warps to the first general-purpose warp scheduling and instruction issuing unit 230_1. The first general-purpose warp scheduling and instruction issuing unit 230_1 then issues the non-tensor computation instructions for the current non-tensor warp to the 2_1st general-purpose computing core 260_1 (step S350). This embodiment introduces the first tensor warp scheduling and instruction issuance unit 220_1, dedicated to the 2_1st tensor core 250_1, to decouple the 2_1st tensor core 250_1 from the 2_1st general-purpose computing core 260_1 in the second operating mode. Similarly, the second thread block splitting unit 210 dispatches non-tensor warps to the Nth general-purpose warp scheduling and instruction issuance unit 230_N, and the Nth general-purpose warp scheduling and instruction issuance unit 230_N issues non-tensor computing instructions for the non-tensor warps to the 2_Nth general-purpose computing core 260_N (step S350).
[0035] The second shared memory area 270 is coupled to the 2_1st tensor core 250_1 to the 2_Nth tensor core 250_N and the 2_1st general purpose computing core 260_1 to the 2_Nth general purpose computing core 260_N. The 2_1st tensor core 250_1 to the 2_Nth tensor core 250_N and the 2_1st general purpose computing core 260_1 to the 2_Nth general purpose computing core 260_N can exchange data in the second shared memory area 270. The second shared memory area 270 can refer to Figure 1 The related description of the first shared memory area 160 is similar and will not be repeated here.
[0036] In summary, the second thread block splitting unit 210 can dispatch tensor warps involving tensor computation to the first tensor warp scheduling and instruction issuing unit 220_1 to the Nth tensor warp scheduling and instruction issuing unit 220_N, and dispatch non-tensor warps not involving tensor computation to the first general warp scheduling and instruction issuing unit 230_1 to the Nth general warp scheduling and instruction issuing unit 230_N. Based on this, while the 1st to Nth tensor warp scheduling and instruction issuing units 220_1 through 220_N issue tensor computation instructions for tensor warps to the 2nd to 1st tensor cores 250_1 through 2nd to 2nd Nth tensor cores 250_N, the 1st to 2nd general-purpose warp scheduling and instruction issuing units 230_1 through 230_N can simultaneously issue non-tensor computation instructions for non-tensor warps to the 2nd to 1st general-purpose computing cores 260_1 through 2nd to 2nd Nth general-purpose computing cores 260_N. In other words, the 2nd to 2nd tensor cores 250_1 through 2nd to 2nd Nth tensor cores 250_N and the 2nd to 1st general-purpose computing cores 260_1 through 2nd to 2nd Nth general-purpose computing cores 260_N can simultaneously execute their respective computational operations. Therefore, in the second operating mode, the execution efficiency of the second artificial intelligence chip 200 can be effectively improved.
[0037] Figure 4 FIG. 1 is a schematic diagram of a circuit module of a tensor warp scheduling and instruction issuing unit according to an embodiment of the present invention. Figure 4 The tensor warp scheduling and instruction issuing unit 220 shown may serve as Figure 2 The illustrated example is one of many implementations of each of the 1st tensor warp scheduling and instruction issuing unit 220_1 to the Nth tensor warp scheduling and instruction issuing unit 220_N. Figure 4 The second thread block splitting unit 210, the tensor warp scheduling and instruction issuing unit 220, the second warp synchronization unit 240 and the tensor core 250 can be referred to in Figure 1 Relevant descriptions of the second thread block splitting unit 210 , the first tensor warp scheduling and instruction issuing unit 220_1 to the Nth tensor warp scheduling and instruction issuing unit 220_N, the second warp synchronization unit 240 , and the 1_1th tensor core 140_1 to the 1_Nth tensor core 140_N are omitted for clarity.
[0038] exist Figure 4In this embodiment, the tensor warp scheduling and instruction issuance unit 220 includes a buffer 221, a warp scheduling unit 222, an instruction decoding unit 223, an operand reading unit 224, a scalar instruction execution unit 225, a scalar register file 226, and a tensor core instruction forwarding unit 227. The buffer 221 is coupled to the second thread block partitioning unit 210. The second thread block partitioning unit 210 dispatches the tensor warp corresponding to the tensor warp scheduling and instruction issuance unit 220 to the buffer 221. The tensor warp includes functions such as reading instructions and checking whether warp synchronization is complete, which will not be described in detail here.
[0039] The warp scheduling unit 222 is coupled to the buffer 221. The warp scheduling unit 222 selects and issues an instruction from a warp in the buffer 221 to the instruction decode unit 223. For example, the warp scheduling unit 222 selects a warp from the warps in the buffer 221 using a round-robin or fixed priority method and issues an instruction from the selected warp to the instruction decode unit 223.
[0040] The instruction decode unit 223 is coupled to the warp scheduling unit 222 to receive instructions. The instruction decode unit 223 decodes the received instruction and generates a decoded result for the operand fetch unit 224. For example, the instruction decode unit 223 parses the instruction field and passes the parsed instruction opcode, operand address, and other information to the operand fetch unit 224. The operand fetch unit 224 is coupled to the instruction decode unit 223, the scalar register file 226, the scalar instruction execution unit 225, and the tensor core instruction forwarding unit 227.
[0041] The instruction segments belonging to the tensor warp in a program contain not only tensor core instructions but also some scalar computation instructions. These scalar computation instructions are used to calculate information such as the starting address and matrix size for subsequent tensor core instructions. Therefore, a scalar register file 226 and a scalar instruction execution unit 225 are required. In response to the decoding result of the instruction decoding unit 223 indicating a scalar computation instruction, the operand reading unit 224 sends the scalar computation instruction to the scalar instruction execution unit 225. The scalar instruction execution unit 225 writes the execution result of the scalar computation instruction (e.g., starting address, matrix size, and other information) to the scalar register file 226.
[0042] In response to the decoded result from the instruction decode unit 223 indicating a tensor core instruction, the operand read unit 224 reads the corresponding information of the tensor core instruction (e.g., starting address, matrix size, etc.) from the scalar register file 226 based on the decoded result. The operand read unit 224 sends the tensor core instruction and the corresponding information to the tensor core instruction forwarding unit 227. The tensor core instruction forwarding unit 227 sends the tensor core instruction and the corresponding information to the corresponding tensor core 250.
[0043] The advantages of this embodiment include: 1. General-purpose computing cores and tensor cores have their own thread warp scheduling and instruction issuance mechanisms. The general-purpose computing cores and tensor cores can execute instructions completely asynchronously without blocking each other.
[0044] 2. If the tensor core finishes all calculations first, the thread warp can exit in time and the tensor core starts the calculation of the next thread warp, eliminating the long tail effect and improving the utilization of the tensor core computing power.
[0045] 3. The thread bundles of tensor cores and general-purpose computing cores logically belong to the same thread block. Synchronization between tensor cores and general-purpose computing cores is still completed within the same thread block, making intra-thread block synchronization and data exchange very convenient.
[0046] 4. Backward compatibility. If the switch for independent warp scheduling and instruction issuance for tensor cores (independent scheduling switch register 211) is disabled, that is, if the second thread block splitting unit 210 operates in the first operating mode, the second artificial intelligence chip 200 is fully compatible with solutions where tensor cores and general-purpose computing cores share warps.
[0047] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An artificial intelligence chip, characterized in that: The artificial intelligence chip includes: Multiple Tensor Cores; a plurality of tensor warp scheduling and instruction-issuing units, wherein a first tensor warp scheduling and instruction-issuing unit of the plurality of tensor warp scheduling and instruction-issuing units is coupled to a first tensor core of the plurality of tensor cores; Multiple general-purpose computing cores; a plurality of general-purpose warp scheduling and instruction issuing units, wherein a first general-purpose warp scheduling and instruction issuing unit among the plurality of general-purpose warp scheduling and instruction issuing units is coupled to the first tensor core and a first general-purpose computing core among the plurality of general-purpose computing cores; and a thread block splitting unit coupled to the plurality of tensor warp scheduling and instruction issuing units and the plurality of general-purpose warp scheduling and instruction issuing units, wherein the thread block splitting unit splits a current thread block into a plurality of warps, In response to the thread block slicing unit operating in a first operating mode, the thread block slicing unit dispatches each of the plurality of thread warps to one of the plurality of general-purpose thread warp scheduling and instruction issuing units, and the first general-purpose thread warp scheduling and instruction issuing unit issues each instruction of the current thread warp to one of the first tensor core and the first general-purpose computing core; and in response to the thread block slicing unit operating in a second operating mode, the thread block slicing unit dispatches each of at least one tensor thread warp involved in tensor computation among the plurality of thread warps to one of the plurality of tensor thread warp scheduling and instruction issuing units, the first tensor thread warp scheduling and instruction issuing unit issues the tensor computation instruction of the current tensor thread warp to the first tensor core, the thread block slicing unit dispatches each of at least one non-tensor thread warp not involved in the tensor computation among the plurality of thread warps to one of the plurality of general-purpose thread warp scheduling and instruction issuing units, and the first general-purpose thread warp scheduling and instruction issuing unit issues the non-tensor computation instruction of the current non-tensor thread warp to the first general-purpose computing core.
2. The artificial intelligence chip according to claim 1, characterized in that: A second tensor warp scheduling and instruction issuing unit among the plurality of tensor warp scheduling and instruction issuing units is coupled to a second tensor core among the plurality of tensor cores, and a second general purpose warp scheduling and instruction issuing unit among the plurality of general purpose warp scheduling and instruction issuing units is coupled to the second tensor core and a second general purpose computing core among the plurality of general purpose computing cores. In response to the thread block splitting unit operating in the first operation mode, the second general warp scheduling and instruction issuing unit issues each instruction of the local current warp to one of the second tensor core and the second general computing core; In response to the thread block splitting unit operating in the second operation mode, the second tensor warp scheduling and instruction issuing unit issues the tensor computing instructions of the local current tensor warp to the second tensor core, and the second general-purpose warp scheduling and instruction issuing unit issues the non-tensor computing instructions of the local current non-tensor warp to the second general-purpose computing core.
3. The artificial intelligence chip according to claim 1, characterized in that: The thread block segmentation unit includes: an independent scheduling switch register, wherein a state of the independent scheduling switch register depends on a program descriptor, In response to the independent scheduling switch register indicating that the switch is closed, the thread block slicing unit operates in the first operation mode; and in response to the independent scheduling switch register indicating that the switch is open, the thread block slicing unit operates in the second operation mode.
4. The artificial intelligence chip according to claim 1, characterized in that The thread block slicing unit dispatches each of the at least one tensor warp to one of the plurality of tensor warp scheduling and instruction issuing units, and dispatches each of the at least one non-tensor warp to one of the plurality of general warp scheduling and instruction issuing units based on a program descriptor.
5. The artificial intelligence chip according to claim 1, characterized in that: The artificial intelligence chip further includes: A warp synchronization unit is coupled to the plurality of tensor warp scheduling and instruction issuing units and the plurality of general warp scheduling and instruction issuing units to receive a synchronization instruction, wherein the warp synchronization unit selectively controls instruction issuance of one or more of the plurality of tensor warp scheduling and instruction issuing units and the plurality of general warp scheduling and instruction issuing units based on the synchronization instruction.
6. The artificial intelligence chip according to claim 1, characterized in that The artificial intelligence chip further includes: A shared memory area is coupled to the plurality of tensor cores and the plurality of general computing cores, wherein the plurality of tensor cores and the plurality of general computing cores exchange data in the shared memory area.
7. The artificial intelligence chip according to claim 1, characterized in that: Each of the plurality of tensor warp scheduling and instruction issuing units comprises: a buffer coupled to the thread block slicing unit, wherein the thread block slicing unit dispatches a corresponding one of the at least one tensor warp to the buffer; a warp scheduling unit coupled to the buffer, wherein the warp scheduling unit selects and issues an instruction of a warp in the buffer; an instruction decoding unit, coupled to the warp scheduling unit to receive the one instruction, wherein the instruction decoding unit decodes the one instruction to generate a decoding result; scalar register set; scalar instruction execution unit; A tensor core instruction forwarding unit; and an operand reading unit, coupled to the instruction decoding unit, the scalar register group, the scalar instruction execution unit and the tensor core instruction forwarding unit, wherein In response to the decoding result indicating a scalar computing instruction, the operand reading unit sends the scalar computing instruction to the scalar instruction execution unit, and the scalar instruction execution unit writes the execution result of the scalar computing instruction to the scalar register group; and in response to the decoding result indicating a tensor core instruction, the operand reading unit reads corresponding information of the tensor core instruction from the scalar register group based on the decoding result, the operand reading unit sends the tensor core instruction and the corresponding information to the tensor core instruction forwarding unit, and the tensor core instruction forwarding unit sends the tensor core instruction and the corresponding information to a corresponding tensor core among the multiple tensor cores.
8. A method for operating an artificial intelligence chip, characterized in that: The operation method includes: A thread block splitting unit of the artificial intelligence chip splits a current thread block into a plurality of thread warps, wherein the thread block splitting unit is coupled to a plurality of tensor warp scheduling and instruction issuing units of the artificial intelligence chip and a plurality of general-purpose warp scheduling and instruction issuing units of the artificial intelligence chip, a first tensor warp scheduling and instruction issuing unit of the plurality of tensor warp scheduling and instruction issuing units is coupled to a first tensor core of a plurality of tensor cores of the artificial intelligence chip, and a first general-purpose warp scheduling and instruction issuing unit of the plurality of general-purpose warp scheduling and instruction issuing units is coupled to the first tensor core and a first general-purpose computing core of a plurality of general-purpose computing cores of the artificial intelligence chip; In response to the thread block slicing unit operating in a first operating mode, the thread block slicing unit dispatches each of the multiple warps to one of the multiple general-purpose warp scheduling and instruction issuing units, and the first general-purpose warp scheduling and instruction issuing unit issues each instruction of the current warp to one of the first tensor core and the first general-purpose computing core; and in response to the thread block slicing unit operating in a second operating mode, the thread block slicing unit dispatches each of at least one tensor warp involved in tensor computation among the multiple warps to one of the multiple tensor warp scheduling and instruction issuing units, the first tensor warp scheduling and instruction issuing unit issues the tensor computation instruction of the current tensor warp to the first tensor core, the thread block slicing unit dispatches each of at least one non-tensor warp not involved in tensor computation among the multiple warps to one of the multiple general-purpose warp scheduling and instruction issuing units, and the first general-purpose warp scheduling and instruction issuing unit issues the non-tensor computation instruction of the current non-tensor warp to the first general-purpose computing core.
9. The operating method according to claim 8, characterized in that: A second tensor warp scheduling and instruction issuing unit among the plurality of tensor warp scheduling and instruction issuing units is coupled to a second tensor core among the plurality of tensor cores, a second general purpose warp scheduling and instruction issuing unit among the plurality of general purpose warp scheduling and instruction issuing units is coupled to the second tensor core and a second general purpose computing core among the plurality of general purpose computing cores, and the operating method further includes: In response to the thread block splitting unit operating in the first operating mode, the second general-purpose thread warp scheduling and instruction issuing unit issues each instruction of the local current thread warp to the second tensor core and one of the second general-purpose computing cores; and in response to the thread block splitting unit operating in the second operating mode, the second tensor thread warp scheduling and instruction issuing unit issues the tensor computing instructions of the local current tensor thread warp to the second tensor core, and the second general-purpose thread warp scheduling and instruction issuing unit issues the non-tensor computing instructions of the local current non-tensor thread warp to the second general-purpose computing core.
10. The operating method according to claim 8, characterized in that: The operation method further includes: In response to an independent scheduling switch register of the thread block slicing unit indicating that a switch is closed, the thread block slicing unit operates in the first operation mode, wherein a state of the independent scheduling switch register depends on a program descriptor; and in response to the independent scheduling switch register indicating that a switch is open, the thread block slicing unit operates in the second operation mode.
11. The operating method according to claim 8, characterized in that: The operation method further includes: The thread block slicing unit dispatches each of the at least one tensor warp to one of the plurality of tensor warp scheduling and instruction issuing units based on a program descriptor, and dispatches each of the at least one non-tensor warp to one of the plurality of general warp scheduling and instruction issuing units.
12. The operating method according to claim 8, characterized in that: The operation method further includes: The thread warp synchronization unit of the artificial intelligence chip selectively controls the instruction issuance of one or more of the multiple tensor thread warp scheduling and instruction issuance units and the multiple general thread warp scheduling and instruction issuance units based on a synchronization instruction, wherein the thread warp synchronization unit is coupled to the multiple tensor thread warp scheduling and instruction issuance units and the multiple general thread warp scheduling and instruction issuance units to receive the synchronization instruction.
13. The operating method according to claim 8, characterized in that: The operation method further includes: The plurality of tensor cores and the plurality of general computing cores exchange data in a shared memory area of the artificial intelligence chip, wherein the shared memory area is coupled to the plurality of tensor cores and the plurality of general computing cores.
14. The operating method according to claim 8, characterized in that: Each of the plurality of tensor warp scheduling and instruction issuing units includes a buffer, a warp scheduling unit, an instruction decoding unit, an operand reading unit, a scalar register group, a scalar instruction execution unit, and a tensor core instruction forwarding unit. The operating method further includes: dispatching, by the thread block slicing unit, corresponding ones of the at least one tensor warp to the buffer, wherein the buffer is coupled to the thread block slicing unit, and the warp scheduling unit is coupled to the buffer; An instruction of one warp in the buffer is selected and issued by the warp scheduling unit, wherein the instruction decoding unit is coupled to the warp scheduling unit to receive the one instruction; The instruction decoding unit decodes the one instruction to generate a decoding result, wherein the operand reading unit is coupled to the instruction decoding unit, the scalar register group, the scalar instruction execution unit, and the tensor core instruction forwarding unit; In response to the decoding result indicating a scalar computing instruction, the operand read unit sends the scalar computing instruction to the scalar instruction execution unit, and the scalar instruction execution unit writes the execution result of the scalar computing instruction to the scalar register group; and in response to the decoding result indicating a tensor core instruction, the operand read unit reads corresponding information of the tensor core instruction from the scalar register group based on the decoding result, sends the tensor core instruction and the corresponding information to the tensor core instruction forwarding unit, and the tensor core instruction forwarding unit sends the tensor core instruction and the corresponding information to a corresponding tensor core among the multiple tensor cores.
Citation Information
Patent Citations
Artificial intelligence chip, operation method thereof and machine readable storage medium
CN117852600A
Implementation device and method for sharing register block by general-purpose computing core and tensor core
CN120295670A
Vector kernel module of artificial intelligence chip and operation method thereof
CN120469721A
Object supply system and control method thereof
KR1020250064426A
Cited By
Artificial intelligence chip and operation method thereof
CN121116908A
Universal graphics processor, computing method, computing device, medium, and program product
CN121120361A
Processor, chip, network device and wireless communication data processing method
CN121210390A