Method, Processor, and Storage Medium for Performing Matrix Multiplication and Accumulation Operations
By designing a special processor data path, calculating the dot product of vector pairs associated with matrix operation objects and accumulating part of the product, the problem of low efficiency of matrix product and accumulation operations in the prior art is solved, and more efficient matrix operation processing capabilities are achieved.
Patent Information
- Application Number
- CN202111061786.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-11-29
- Filing Date
- 2018-05-07
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2038-05-07
AI Technical Summary
Existing processors are inefficient in performing matrix product and accumulation operations, mainly due to the lack of specialized hardware support, resulting in inefficient use of the input bandwidth of the data path.
A processor data path is designed to generate each element of the result matrix by calculating the dot product of vector pairs associated with the matrix operation object, and to accumulate multiple parts into the result queue using an adder.
This method significantly improves the efficiency of matrix product and accumulation operations, reduces the bandwidth between register files and data path input, and improves the processor's processing capability of matrix operations.
Smart Images

Figure CN113961872B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese Patent Application No. 201810425869.9, filed on May 7, 2018.
[0002] Cross - Reference to Related Applications
[0003] This application claims the benefit of U.S. Provisional Application No. 62 / 503,159, filed on May 8, 2017, entitled "Generalized Acceleration of Matrix Multiply Accumulate Operations" (Attorney Docket No. NVIDP1157+), the entire content of which is incorporated herein by reference. Technical Field
[0004] The present disclosure relates to implementing arithmetic operations on a processor, and more particularly to accelerating matrix multiply accumulate operations. Background Art
[0005] Modern computer processors are basically integrated circuits designed to perform logical tasks. One task that processors are really good at is performing arithmetic operations on numbers encoded in different formats (e.g., 8-bit integers, 32-bit integers, 32-bit floating-point values, etc.). However, most processors contain logic for performing these arithmetic operations on scalar operands. For example, the logic designed to perform an addition operation is designed to perform the operation using two different operands, each operand encoding a specific value to be added to the other operand. However, arithmetic operations are not limited to scalar values. In fact, many applications may use arithmetic operations on vector or matrix inputs. An example of performing an arithmetic operation on a vector is the dot product operation. Although it is common to calculate the dot product in these applications (e.g., physics), modern processors are generally not designed with hardware in the circuit to efficiently perform these operations. Instead, scalar values are used to reduce the higher-level operations to a series of basic arithmetic operations. For example, in a dot product operation, each vector operand includes multiple elements, and the dot product operation is performed by multiplying the corresponding element pairs of the two input vectors to generate multiple partial products (i.e., intermediate results) and then summing the multiple partial products. Each basic arithmetic operation can be sequentially performed using the hardware logic designed into the processor, and the intermediate results can be stored in temporary memory and reused as an operand for another subsequent arithmetic operation.
[0006] Conventional processors include one or more cores, where each core may include an arithmetic logic unit (ALU) and / or a floating-point unit for performing basic operations on integer and / or floating-point values. Conventional floating-point units can be designed to implement fused multiply accumulate (FMA) operations, which multiply two scalar operands and add the intermediate result and an optional third scalar operand to an accumulation register. Matrix multiply and accumulate (MMA) operations are an extension of FMA operations to scalar values applied to matrix operands. In other words, an MMA operation multiplies two matrices and optionally adds the resulting intermediate matrix to a third matrix operand. Fundamentally, an MMA operation can be reduced to a number of basic dot product operations added to an accumulation register. Additionally, a dot product operation can be further reduced to a series of FMA operations on pairs of scalar operands.
[0007] Conventional processors can implement matrix operations by decomposing an MMA operation into a series of dot product operations and addition operations, and each dot product operation can be further decomposed into a series of FMA instructions for corresponding elements of a pair of vectors. However, since an MMA operation must be decomposed into each basic arithmetic operation using scalar operands, this technique is not efficient. Each basic arithmetic operation executed by the logic of the processor involves moving scalar operands between the register file of the processor and the input to the data path (i.e., the logic circuit). However, the fundamental concept of matrix operations is that the same elements of a matrix are reused in multiple dot product operations (e.g., the same row of a first matrix is used to generate multiple dot products corresponding to multiple columns of a second matrix). If each basic arithmetic operation requires loading data from the register file to the input of the data path before performing the arithmetic operation, then each data element of the input operands can be loaded from the register file to the data path many times, which is an inefficient use of the register file bandwidth. Although there may be techniques to improve the efficiency of the processor (e.g., a register file with multiple banks such that operands can be efficiently stored in separate banks and multiple operands can be loaded from the register file to the input of the data path in a single clock cycle), generally, the data path is not designed specifically for matrix operations. Therefore, there is a need to address these problems and / or other problems related to the prior art. Summary of the Invention
[0008] A method, computer-readable medium, and processor for performing matrix multiply-accumulate (MMA) operations are disclosed. The processor includes a data path configured to perform MMA operations to generate a plurality of elements of a result matrix at an output of the data path. Each element of the result matrix is generated by computing at least one dot product of corresponding vector pairs associated with matrix operands specified in an instruction for the MMA operation. The dot product operation includes the steps of: generating a plurality of partial products by multiplying each element of a first vector by a corresponding element of a second vector; aligning the plurality of partial products based on exponents associated with each element of the first vector and each element of the second vector; and adding the plurality of aligned partial products to a result queue using at least one adder. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 FIG. 6 shows a flowchart of a method for performing matrix multiply-accumulate operations according to one embodiment;
[0010] Figure 2 FIG. 7 shows a parallel processing unit (PPU) according to one embodiment;
[0011] Figure 3A FIG. 8 shows a general processing cluster of a PPU according to one embodiment Figure 2 ;
[0012] Figure 3B FIG. 9 shows a partitioning unit of a PPU according to one embodiment Figure 2 ;
[0013] Figure 4 FIG. 10 shows a streaming multiprocessor according to one embodiment Figure 3A ;
[0014] Figure 5 FIG. 11 shows a system-on-chip including a PPU according to one embodiment Figure 2 ;
[0015] Figure 6 FIG. 12 is a conceptual diagram of a graphics processing pipeline implemented by a PPU according to one embodiment Figure 2 ;
[0016] Figure 7 FIG. 13 shows a matrix multiply-accumulate operation according to one embodiment;
[0017] Figure 8 FIG. 14 is a conceptual diagram of a dot product operation according to one embodiment;
[0018] Figure 9 FIG. 15 shows a portion of a processor including a data path configured to implement matrix operations according to one embodiment;
[0019] Figure 10 Shows a conventional double - precision floating - point fused multiply - add data path according to one embodiment;
[0020] Figure 11 Shows a half - precision matrix multiply and accumulate data path according to one embodiment;
[0021] Figure 12 Shows a half - precision matrix multiply and accumulate data path according to another embodiment;
[0022] Figure 13 Shows a half - precision matrix multiply and accumulate data path according to yet another embodiment;
[0023] Figure 14 Shows a half - precision matrix multiply and accumulate data path according to one embodiment, configured to share at least one pipeline stage with Figure 10 the double - precision floating - point fused multiply - add data path of Figure 13 ; and
[0024] Figure 15 Shows an exemplary system in which various architectures and / or functions of the various previous embodiments can be implemented. Detailed Description
[0025] Many modern applications can benefit from more efficient processing of matrix operations by a processor. Arithmetic operations performed on matrix operation objects are commonly used in various algorithms, including but not limited to: deep learning algorithms, linear algebra, and graphics acceleration, etc. Higher efficiency can be obtained by using parallel processing units because matrix operations can be reduced to multiple parallel operations on different parts of the matrix operation objects.
[0026] A new paradigm for data - path design is explored herein to accelerate matrix operations as performed by a processor. The basic concept of the data path is that the data path performs one or more dot - product operations on multiple vector operation objects. Matrix operations can then be accelerated by reducing them to multiple dot - product operations, and some dot - product operations can benefit from data sharing within the data path, which reduces the bandwidth between the register file and the input of the data path.
[0027] Figure 1FIG. 0 shows a flowchart of a method 100 for performing matrix multiply and accumulate operations according to one embodiment. It will be appreciated that method 100 is described within the scope of software executed by a processor; however, in some embodiments, method 100 may be implemented in hardware or in some combination of hardware and software. Method 100 begins at step 102, where an instruction for a matrix multiply and accumulate (MMA) operation is received. In one embodiment, the instruction for the MMA operation specifies a plurality of matrix operands. The first operand specifies a multiplicand input matrix A, the second operand specifies a multiplier input matrix B, and the third operand specifies an accumulator matrix C for accumulating the product results of the first two input matrices. Each operand specified in the instruction is a matrix of a plurality of elements in a two-dimensional array having rows and columns.
[0028] At step 104, at least two vectors of the first operand specified in the instruction and at least two vectors of the second operand specified in the instruction are loaded from a register file into a plurality of operand collectors. In one embodiment, the operand collectors are a plurality of flip-flops coupled to an input of a data path configured to perform an MMA operation. The plurality of flip-flops temporarily store data of the operands for the MMA instruction at the input of the data path such that the plurality of operands can be loaded from the register file to the input of the data path over a plurality of clock cycles. Generally, the register file has a limited amount of bandwidth on one or more read ports such that only a limited amount of data can be read from the register file in a given clock cycle. Thus, the operand collectors enable all of the operands required for the data path to be read from the data file over a plurality of clock cycles before initiating the execution of the MMA operation on the data path.
[0029] In step 106, an MMA operation is performed to generate multiple elements of the result matrix at the output of the data path. In one embodiment, each element of the result matrix is generated by computing at least one dot product of corresponding vector pairs stored in multiple operand collectors. The data path can be designed to generate multiple elements of the result matrix in multiple passes of the data path while consuming different combinations of vectors stored in the operand collectors during each pass. Optionally, using different sets of logic to compute multiple dot products in parallel, the data path can be designed to generate multiple elements of the result matrix in a single pass of the data path. Of course, in some embodiments, to generate more result matrix elements in a single instruction cycle, multiple sets of logic can be used to compute multiple dot products in parallel and multiple passes of the data path can be used. It should be appreciated that in subsequent passes or instruction cycles, multiple elements of the result matrix are generated without the need to load new operand data from the register file into the operand collectors. Further, it should be appreciated that each vector of the input matrix operands (i.e., A and B) stored in the operand collectors can be consumed by multiple dot product operations for multiple elements of the result matrix.
[0030] In accordance with the user's requirements, more illustrative information regarding various alternative architectures and features that can or cannot implement the foregoing framework will now be presented. It should be noted specifically that the following information is presented for illustrative purposes and should not be construed as limiting in any way. Any of the following features can optionally be incorporated or excluded in combination with other features described.
[0031] Parallel processing architecture
[0032] Figure 2FIG. 200 shows a parallel processing unit (PPU) 200 according to one embodiment. In one embodiment, the PPU 200 is a multi-threaded processor implemented on one or more integrated circuit devices. The PPU 200 is a potential hidden architecture designed to parallel process a large number of threads. A thread (i.e., an execution thread) is an instance of a set of instructions configured to be executed by the PPU 200. In one embodiment, the PPU 200 is a graphics processing unit (GPU) configured to implement a graphics rendering pipeline for processing three-dimensional (3D) graphics data to generate two-dimensional (2D) image data for display on a display device (such as a liquid crystal display (LCD) device). In other embodiments, the PPU 200 may be used to perform general computing. Although an exemplary parallel processor is provided herein for illustrative purposes, it should be noted specifically that such a processor is set forth for illustrative purposes only, and any processor may be used to supplement and / or replace the processor.
[0033] As Figure 2 shown, the PPU 200 includes an input / output (I / O) unit 205, a host interface unit 210, a front-end unit 215, a scheduler unit 220, a work distribution unit 225, a hub 230, a crossbar (Xbar) 270, one or more general processing clusters (GPCs) 250, and one or more partition units 280. The PPU 200 may be connected to a host processor or other peripheral devices via a system bus 202. The PPU 200 may also be connected to a local memory including a plurality of storage devices 204. In one embodiment, the local memory may include a plurality of dynamic random access memory (DRAM) devices.
[0034] The I / O unit 205 is configured to send and receive communications (i.e., commands, data, etc.) from a host processor (not shown) via the system bus 202. The I / O unit 205 may communicate directly with the host processor via the system bus 202 or through one or more intermediate devices (such as a memory bridge). In one embodiment, the I / O unit 205 implements a Peripheral Component Interconnect Express (PCIe) interface for communication via a PCIe bus. In alternative embodiments, the I / O unit 205 may implement other types of known interfaces for communicating with external devices.
[0035] The I / O unit 205 is coupled to the host interface unit 210, which decodes data packets received via the system bus 202. In one embodiment, the data packets represent commands configured to cause the PPU 200 to perform various operations. The host interface unit 210 sends the decoded commands to various other units of the PPU 200 in the manner specified by the commands. For example, some commands can be sent to the front-end unit 215. Other commands can be sent to the hub 230 or other units of the PPU 200, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). In other words, the host interface unit 210 is configured to route communications between the various logical units of the PPU 200.
[0036] In one embodiment, a program executed by the host processor encodes a command stream in a buffer, which provides the workload for the PPU 200 to process. The workload can include a large number of instructions and data to be processed by those instructions. The buffer is an area in memory that can be accessed (i.e., read / written) by both the host processor and the PPU 200. For example, the host interface unit 210 can be configured to access the buffer in the system memory connected to the system bus 202 via a memory request sent by the I / O unit 205 on the system bus 202. In one embodiment, the host processor writes the command stream to the buffer and then sends a pointer to the start of the command stream to the PPU 200. The host interface unit 210 provides pointers to one or more command streams to the front-end unit 215. The front-end unit 215 manages one or more streams, reads commands from the streams, and forwards the commands to the various units of the PPU 200.
[0037] The front-end unit 215 is coupled to the scheduler unit 220, which configures the various GPCs 250 to process the tasks defined by one or more streams. The scheduler unit 220 is configured to track status information related to the various tasks managed by the scheduler unit 220. The status can indicate which GPC 250 the task is assigned to, whether the task is active or inactive, the priority associated with the task, etc. The scheduler unit 220 manages the execution of multiple tasks on one or more GPCs 250.
[0038] The scheduler unit 220 is coupled to a work distribution unit 225, which is configured to dispatch tasks to be executed on the GPCs 250. The work distribution unit 225 may keep track of a plurality of scheduled tasks received from the scheduler unit 220. In one embodiment, the work distribution unit 225 manages a pending task pool and an active task pool for each of the GPCs 250. The pending task pool may include a plurality of time slots (e.g., 32 time slots) that contain tasks assigned to be processed by a particular GPC 250. The active task pool may include a plurality of time slots (e.g., 4 time slots) for tasks that the GPC 250 is actively processing. When a GPC 250 finishes executing a task, the task is evicted from the active task pool of the GPC 250, and one of the other tasks from the pending task pool is selected and scheduled to be executed on the GPC 250. If an active task on a GPC 250 has become idle, e.g., while waiting for a data dependency to be resolved, then the active task may be evicted from the GPC 250 and returned to the pending task pool, while another task from the pending task pool is selected and scheduled to be executed on the GPC 250.
[0039] The work distribution unit 225 communicates with one or more GPCs 250 via an XBar 270. The XBar 270 is an interconnect network that couples many of the units of the PPU 200 to other units of the PPU 200. For example, the XBar 270 may be configured to couple the work distribution unit 225 to a particular GPC 250. Although not explicitly shown, one or more other units of the PPU 200 are coupled to the host unit 210. The other units may also be connected to the XBar 270 via a hub 230.
[0040] Tasks are managed by the scheduler unit 220 and dispatched by the work distribution unit 225 to the GPCs 250. The GPCs 250 are configured to process the tasks and generate results. The results may be consumed by other tasks within the GPC 250, routed via the XBar 270 to a different GPC 250, or stored in the memory 204. The results may be written to the memory 204 via a partitioning unit 280, which implements a memory interface for writing to / reading from the memory 204. In one embodiment, the PPU 200 includes a number U of partitioning units 280 that is equal to the number of separate and distinct storage devices 204 coupled to the PPU 200. The partitioning unit 280 will be described in more detail below in connection with Figure 3B will be described in more detail.
[0041] In one embodiment, the host processor executes a driver kernel that implements an application programming interface (API) that enables one or more applications executed on the host processor to schedule operations to be performed on the PPU 200. The application can generate instructions (i.e., API calls) that cause the driver kernel to generate one or more tasks to be performed by the PPU 200. The driver kernel outputs the tasks to one or more streams processed by the PPU 200. Each task can include one or more related thread groups, referred to herein as warps. A thread block can refer to a plurality of thread groups that include instructions for performing a task. Threads in the same thread group can exchange data through a shared memory. In one embodiment, a thread group includes 32 related threads.
[0042] Figure 3A According to an embodiment Figure 2 PPU 200 GPC 250. Figure 3A As shown, each GPC 250 includes multiple hardware units for processing tasks. In one embodiment, each GPC 250 includes a pipeline manager 310, a pre-raster operation (PROP) unit 315, a raster engine 325, a work distribution crossbar (WDX) 380, a memory management unit (MMU) 390, and one or more texture processing clusters (TPC) 320. It should be appreciated that Figure 3A The GPC 250 may include instead Figure 3A The units shown in or except Figure 3A Other hardware units other than those shown in .
[0043] In one embodiment, the operation of the GPC 250 is controlled by the pipeline manager 310. The pipeline manager 310 manages the configuration of one or more TPCs 320 for processing tasks assigned to the GPC 250. In one embodiment, the pipeline manager 310 may configure at least one of the one or more TPCs 320 to implement at least a portion of the graphics rendering pipeline. For example, the TPC 320 may be configured to execute a vertex shader program on the programmable streaming multiprocessor (SM) 340. The pipeline manager 310 may also be configured to route data packets received from the work distribution unit 225 to appropriate logic units within the GPC 250. For example, some data packets may be routed to the fixed function hardware units in the PROP 315 and / or raster engine 325, while other data packets may be routed to the TPC 320 for processing by the primitive engine 335 or SM 340.
[0044] The PROP unit 315 is configured to route data generated by the raster engine 325 and TPC 320 to the Raster Operation (ROP) unit in the partition unit 280, which will be described in more detail below. The PROP unit 315 may also be configured to perform optimizations for color blending, organize pixel data, perform address translation, and the like.
[0045] The raster engine 325 includes a plurality of fixed function hardware units configured to perform various raster operations. In one embodiment, the raster engine 325 includes a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, and a tile stitching engine. The setup engine receives the transformed vertices and generates plane equations associated with the geometric primitives defined by the vertices. The plane equations are transmitted to the coarse raster engine to generate coverage information for the primitives (e.g., x, y coverage masks for tiles). The output of the coarse raster engine may be transmitted to the culling engine, where fragments associated with primitives that fail the z-test are culled, and to the clipping engine, where fragments located outside the view frustum are clipped off. Those fragments that are spared from clipping and culling may be passed to the fine raster engine to generate the attributes of the pixel fragments based on the plane equations generated by the setup engine. The output of the raster engine 325 includes the fragments to be processed, such as fragment shading implemented within the TPC 320.
[0046] Each TPC 320 included in the GPC 250 includes an M-Pipe Controller (MPC) 330, a primitive engine 335, one or more SMs 340, and one or more texture units 345. The MPC 330 controls the operation of the TPC 320 and routes data packets received from the pipeline manager 310 to the appropriate units within the TPC 320. For example, data packets associated with vertices can be routed to the primitive engine 335, which is configured to obtain vertex attributes associated with the vertices from the memory 204. Conversely, data packets associated with shader programs can be sent to the SM 340.
[0047] In one embodiment, the texture unit 345 is configured to load a texture map (e.g., a 2D array of texture pixels) from the memory 204 and sample the texture map to produce sampled texture values for use by shader programs executed by the SM 340. The texture unit 345 implements texture operations such as filtering operations using mip mapping (i.e., texture maps at different levels of detail). The texture unit 345 also serves as the load / store path from the SM 340 to the MMU 390. In one embodiment, each TPC 320 includes two (2) texture units 345.
[0048] The SM 340 includes programmable streaming processors that are configured to process tasks represented by multiple threads. Each SM 340 is multi-threaded and is configured to execute multiple threads (e.g., 32 threads) from a particular thread group simultaneously. In one embodiment, the SM 340 implements a SIMD (Single Instruction, Multiple Data) architecture, where each thread in the thread group (i.e., warp) is configured to process different data sets based on the same instruction set. All threads in the thread group execute the same instruction. In another embodiment, the SM 340 implements a SIMT (Single Instruction, Multiple Threads) architecture, where each thread in the thread group is configured to process different data sets based on the same instruction set, but where individual threads in the thread group are allowed to diverge during execution. In other words, when an instruction for the thread group is dispatched for execution, some threads in the thread group can be active and thus execute the instruction, while other threads in the thread group can be inactive and thus execute a no-operation (NOP) instead of the instruction. This is described in more detail below in conjunction with Figure 4 a more detailed description of the SM 340.
[0049] The MMU 390 provides an interface between the GPC 250 and the partition unit 280. The MMU 390 can provide virtual address to physical address translation, memory protection, and arbitration of memory requests. In one embodiment, the MMU 390 provides one or more translation lookaside buffers (TLBs) for improving the translation of virtual addresses to physical addresses in the memory 204.
[0050] Figure 3B illustrates a Figure 2 partition unit 280 of the PPU 200 according to one embodiment. As Figure 3B shown, the partition unit 280 includes a raster operation (ROP) unit 350, a secondary (L2) cache 360, a memory interface 370, and an L2 crossbar (XBar) 365. The memory interface 370 is coupled to the memory 204. The memory interface 370 can implement 16-, 32-, 64-, 128-bit data buses, etc. for high-speed data transfer. In one embodiment, the PPU 200 includes U memory interfaces 370, with one memory interface 370 for each partition unit 280, where each partition unit 280 is connected to a corresponding storage device 204. For example, the PPU 200 can be connected to up to U storage devices 204, such as graphics double-data-rate, version 5, synchronous dynamic random access memory (GDDR5 SDRAM). In one embodiment, the memory interface 370 implements a DRAM interface and U equals 8.
[0051] In one embodiment, the PPU 200 implements a multi-level memory hierarchy. The memory 204 is located off-chip in the SDRAM coupled to the PPU 200. Data from the memory 204 can be fetched and stored in the L2 cache 360 located on the chip and shared among the respective GPCs 250. As shown, each partition unit 280 includes a portion of the L2 cache 360 associated with the corresponding storage device 204. Then, lower-level caches can be implemented in the respective units within the GPC 250. For example, each of the SMs 340 can implement a level 1 (L1) cache. The L1 cache is a dedicated memory for a specific SM 340. Data from the L2 cache 360 can be fetched and stored in each of the L1 caches for processing in the functional units of the SM 340. The L2 cache 360 is coupled to the memory interface 370 and the XBar 270.
[0052] The ROP unit 350 includes a ROP manager 355, a color ROP (CROP) unit 352, and a Z ROP (ZROP) unit 354. The CROP unit 352 performs raster operations related to pixel color, such as color compression, pixel blending, and the like. The ZROP unit 354 implements depth testing in conjunction with the raster engine 325. The ZROP unit 354 receives the depth of a sampling position associated with a pixel fragment from a culling engine of the raster engine 325. The ZROP unit 354 tests the depth relative to the corresponding depth of the sampling position associated with the fragment in the depth buffer. If the fragment passes the depth test of the sampling position, the ZROP unit 354 updates the depth buffer and sends the result of the depth test to the raster engine 325. The ROP manager 355 controls the operation of the ROP unit 350. It should be appreciated that the number of partition units 280 may be different than the number of GPCs 250, and thus each ROP unit 350 may be coupled to each of the GPCs 250. Thus, ROP manager 355 tracks packets received from different GPCs 250 and determines to which GPC 250 the results generated by ROP unit 350 are routed. CROP unit 352 and ZROP unit 354 are coupled to L2 cache 360 via L2 XBar 365.
[0053] Figure 4 According to an embodiment Figure 3A The streaming multiprocessor 340. Figure 4 As shown, SM 340 includes an instruction cache 405, one or more scheduler units 410, a register file 420, one or more processing cores 450, one or more special function units (SFU) 452, one or more load / store units (LSU) 454, an interconnection network 480, a shared memory 470, and an L1 cache 490.
[0054] As described above, the work distribution unit 225 dispatches tasks for execution on the GPCs 250 of the PPU 200. The tasks are assigned to specific TPCs 320 within the GPC 250, and if the task is associated with a shader program, the task can be assigned to an SM 340. The scheduler unit 410 receives tasks from the work distribution unit 225 and manages the instruction scheduling for one or more thread groups (i.e., warps) assigned to the SM 340. The scheduler unit 410 schedules threads for execution in parallel thread groups, where each group is called a warp. In one embodiment, each warp includes 32 threads. The scheduler unit 410 can manage multiple different warps, schedule the warps for execution, and then dispatch instructions to the respective functional units (i.e., the cores 450, SFUs 452, and LSUs 454) from multiple different warps during each clock cycle.
[0055] In one embodiment, each scheduler unit 410 includes one or more instruction dispatch units 415. Each dispatch unit 415 is configured to transfer instructions to one or more of the functional units. In Figure 4 the illustrated embodiment, the scheduler unit 410 includes two dispatch units 415, which enables the dispatch of two different instructions from the same warp during each clock cycle. In an alternative embodiment, each scheduler unit 410 may include a single dispatch unit 415 or additional dispatch units 415.
[0056] Each SM 340 includes a register file 420 that provides a set of registers for the functional units of the SM 340. In one embodiment, the register file 420 is partitioned among each of the functional units such that each functional unit is assigned a dedicated portion of the register file 420. In another embodiment, the register file 420 is partitioned among the different warps being executed by the SM 340. The register file 420 provides temporary storage for the operands of the data paths connected to the functional units.
[0057] Each SM 340 includes L processing cores 450. In one embodiment, the SM 340 includes a large number (e.g., 128, etc.) of different processing cores 450. Each core 450 may include a fully pipelined, single-precision processing unit that includes a floating-point arithmetic logic unit and an integer arithmetic logic unit. The core 450 may also include a double-precision processing unit that includes a floating-point arithmetic logic unit. In one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point arithmetic. Each SM 340 also includes M SFUs 452 that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.) and N LSUs 454 that implement load and store operations between the shared memory 470 or L1 cache 490 and the register file 420. In one embodiment, the SM 340 includes 128 cores 450, 32 SFUs 452, and 32 LSUs 454.
[0058] Each SM 340 includes an interconnect network 480 that connects each of the functional units to the register file 420 and connects the LSU 454 to the register file 420, the shared memory 470, and the L1 cache 490. In one embodiment, the interconnect network 480 is a crossbar switch that can be configured to connect any functional unit to any register in the register file 420 and connect the LSU 454 to the register file and memory locations in the shared memory 470 and the L1 cache 490.
[0059] The shared memory 470 is an on-chip memory array that allows data storage and communication between the SM 340 and the primitive engine 335 and between the threads in the SM 340. In one embodiment, the shared memory 470 includes a storage capacity of 64 KB. The L1 cache 490 is located in the path from the SM 340 to the partitioning unit 280. The L1 cache 490 can be used for caching reads and writes. In one embodiment, the L1 cache 490 includes a storage capacity of 24 KB.
[0060] The above-mentioned PPU 200 can be configured to perform highly parallel computing faster than a traditional CPU. Parallel computing has advantages in graphics processing, data compression, biometrics, stream processing algorithms, etc.
[0061] When configured for general-purpose parallel computing, a simpler configuration can be used. In this model, as Figure 2As shown, the fixed-function graphics processing unit is bypassed, creating a simpler programming model. In this configuration, the work distribution unit 225 directly assigns and distributes thread blocks to the TPCs 320. Threads within a block execute the same program, using the unique thread ID in the computation to ensure that each thread generates a unique result, using the SMs 340 to execute the program and perform the computation, using the shared memory 470 to communicate between threads, and using the LSU 454 to read and write to global memory through the partitioned L1 cache 490 and the partitioning unit 280.
[0062] When configured for general-purpose parallel computing, the SMs 340 can also write commands that the scheduler unit 220 can use to start new work on the TPCs 320.
[0063] In one embodiment, the PPU 200 includes a graphics processing unit (GPU). The PPU 200 is configured to receive commands specifying a shader program for processing graphics data. Graphics data can be defined as a set of primitives, such as points, lines, triangles, quads, triangle strips, etc. Typically, a primitive includes data specifying multiple vertices of the primitive (e.g., in a model space coordinate system) and attributes associated with each vertex of the primitive. The PPU 200 can be configured to process the primitives to generate a frame buffer (i.e., pixel data for each of the pixels on a display).
[0064] The application writes the model data of the scene (i.e., a collection of vertices and attributes) to a memory such as system memory or memory 204. The model data defines each of the objects that may be visible on the display. The application then makes an API call to the driver kernel, which requests the model data to be rendered and displayed. The driver kernel reads the model data and writes commands to one or more streams to perform operations to process the model data. These commands can reference different shader programs to be implemented on the SM 340 of the PPU 200, including one or more of vertex shading, hull shading, domain shading, geometry shading, and pixel shading. For example, one or more of the SM 340s can be configured to execute a vertex shader program that processes the multiple vertices defined by the model data. In one embodiment, different SM 340s can be configured to execute different shader programs simultaneously. For example, a first subset of the SM 340s can be configured to execute a vertex shader program, while a second subset of the SM 340s can be configured to execute a pixel shader program. The first subset of the SM 340s processes the vertex data to produce processed vertex data and writes the processed vertex data to the L2 cache 360 and / or memory 204. After the processed vertex data is rasterized (i.e., converted from three-dimensional data into two-dimensional data in screen space) to produce fragment data, the second subset of the SM 340s performs pixel shading to produce processed fragment data, which is then blended with other processed fragment data and written to the frame buffer in memory 204. The vertex shader program and the pixel shader program can be executed simultaneously to process different data from the same scene in a pipelined manner until all the model data of the scene has been rendered to the frame buffer. Then, the content of the frame buffer is transferred to the display controller for display on the display device.
[0065] The PPU 200 can be included in a desktop computer, a laptop computer, a tablet computer, a smart phone (e.g., a wireless, handheld device), a personal digital assistant (PDA), a digital camera, a handheld electronic device, etc. In one embodiment, the PPU 200 is embodied on a single semiconductor substrate. In another embodiment, the PPU 200 is included in a system-on-chip (SoC) together with one or more other logic units such as a reduced instruction set computer (RISC) CPU, a memory management unit (MMU), a digital-to-analog converter (DAC), etc.
[0066] In one embodiment, the PPU 200 may be included on a graphics card, which includes one or more storage devices 204 (such as GDDR5 SDRAM). The graphics card may be configured to interface with a PCIe slot on the motherboard of a desktop computer, which includes, for example, a north bridge chipset and a south bridge chipset. In yet another embodiment, the PPU 200 may be an integrated graphics processing unit (iGPU) included in the chipset (i.e., north bridge) of the motherboard.
[0067] Figure 5 A system-on-chip (SoC) 500 including the Figure 2 PPU 200 is shown according to one embodiment. As Figure 5 shown, as described above, the SoC 500 includes a CPU 550 and a PPU 200. The SoC 500 may also include a system bus 202 to enable communication between the various components of the SoC 500. Memory requests generated by the CPU 550 and the PPU 200 may be routed through a system MMU 590 shared by multiple components of the SoC 500. The SoC 500 may also include a memory interface 595 coupled to one or more storage devices 204. The memory interface 595 may implement, for example, a DRAM interface.
[0068] Although not explicitly shown, in addition to the Figure 5 components shown, the SoC 500 may also include other components. For example, the SoC 500 may include multiple PPUs 200 (e.g., four PPUs 200), video encoders / decoders, and wireless broadband transceivers, among other components. In one embodiment, the SoC 500 may be included in a package-on-package (PoP) configuration together with the memory 204.
[0069] Figure 6 is a conceptual diagram of a graphics processing pipeline 600 implemented by the Figure 2 PPU 200 according to one embodiment. The graphics processing pipeline 600 is an abstract flow chart of processing steps implemented to generate 2D computer-generated images from 3D geometric data. As is well known, a pipeline architecture can execute long-latency operations more efficiently by dividing the operations into multiple stages, where the output of each stage is coupled to the input of the next consecutive stage. Thus, the graphics processing pipeline 600 receives input data 601 that is transferred from one stage of the graphics processing pipeline 600 to the next stage to generate output data 602. In one embodiment, the graphics processing pipeline 600 may represent the one implemented by The graphics processing pipeline defined by the API. Alternatively, the graphics processing pipeline 600 may be implemented in the context of the functionality and architecture of the previous figures and / or any one or more subsequent figures.
[0070] As Figure 6 shown, the graphics processing pipeline 600 includes a pipeline architecture that includes multiple stages. These stages include, but are not limited to, a data assembly stage 610, a vertex shading stage 620, a primitive assembly stage 630, a geometry shading stage 640, a viewport scale, cull, and clip (VSCC) stage 650, a rasterization stage 660, a fragment shading stage 670, and a raster operation stage 680. In one embodiment, the input data 601 includes commands that configure the processing unit to implement the stages of the graphics processing pipeline 600 and configure geometric primitives (e.g., points, lines, triangles, quads, triangle strips, or fans, etc.) to be processed by these stages. The output data 602 may include pixel data (i.e., color data) that is copied into a frame buffer in memory or other type of surface data structure.
[0071] The data assembly stage 610 receives the input data 601, which specifies vertex data for high-order surfaces, primitives, etc. The data assembly stage 610 collects the vertex data in temporary storage or a queue, such as by receiving a command from the host processor that includes a pointer to a buffer in memory and reading the vertex data from that buffer. The vertex data is then transmitted to the vertex shading stage 620 for processing.
[0072] The vertex shading stage 620 processes the vertex data by performing a set of operations (i.e., vertex shader or program) on each of the vertices. The vertices may be specified, for example, as 4-coordinate vectors (i.e., <x, y, z, w>) associated with one or more vertex attributes (e.g., color, texture coordinates, surface normal, etc.). The vertex shading stage 620 may manipulate the individual vertex attributes, such as position, color, texture coordinates, etc. In other words, the vertex shading stage 620 performs operations on the vertex coordinates or other vertex attributes associated with the vertices. These operations typically include lighting operations (i.e., modifying the color attribute of the vertex) and transformation operations (i.e., modifying the coordinate space of the vertex). For example, vertices may be specified using coordinates in an object coordinate space, which are transformed by multiplying the coordinates by a matrix that transforms the coordinates from the object coordinate space to the world space or the normalized-device-coordinate (NCD) space. The vertex shading stage 620 generates the transformed vertex data that is transmitted to the primitive assembly stage 630.
[0073] The primitive assembly stage 630 collects the vertices output by the vertex shading stage 620 and groups the vertices into geometric primitives for processing by the geometry shading stage 640. For example, the primitive assembly stage 630 may be configured to group every three consecutive vertices into a geometric primitive (i.e., a triangle) for transmission to the geometry shading stage 640. In some embodiments, a particular vertex may be reused for consecutive geometric primitives (e.g., two consecutive triangles in a triangle strip may share two vertices). The primitive assembly stage 630 transmits the geometric primitives (i.e., the set of associated vertices) to the geometry shading stage 640.
[0074] The geometry shading stage 640 processes the geometric primitives by performing a set of operations (i.e., a geometry shader or program) on the geometric primitives. Tessellation operations may generate one or more geometric primitives from each geometric primitive. In other words, the geometry shading stage 640 may subdivide each geometric primitive into a finer mesh of two or more geometric primitives for processing by the remainder of the graphics processing pipeline 600. The geometry shading stage 640 transmits the geometric primitives to the viewport SCC stage 650.
[0075] In one embodiment, the graphics processing pipeline 600 may operate within the streaming multiprocessors and the vertex shading stage 620, the primitive assembly stage 630, the geometry shading stage 640, the fragment shading stage 670, and / or the associated hardware / software, and may perform processing operations sequentially. Once the sequential processing operations are complete, in one embodiment, the viewport SCC stage 650 may utilize the data. In one embodiment, the primitive data processed by one or more of the stages in the graphics processing pipeline 600 may be written into a cache (e.g., an L1 cache, a vertex cache, etc.). In such a case, in one embodiment, the viewport SCC stage 650 may access the data in the cache. In one embodiment, the viewport SCC stage 650 and the rasterization stage 660 are implemented as fixed function circuits.
[0076] The viewport SCC stage 650 performs viewport scaling, culling, and clipping of geometric primitives. Each surface being rendered is associated with an abstract camera position. The camera position represents the position of the viewer looking at the scene and defines the viewing frustum of the objects enclosing the scene. The viewing frustum can include viewing planes, a back plane, and four clipping planes. Any geometric primitive that is completely outside the viewing frustum can be culled (i.e., discarded) because these geometric primitives will not contribute to the ultimately rendered scene. Any geometric primitive that is partially inside and partially outside the viewing frustum can be clipped (i.e., transformed into a new geometric primitive enclosed within the viewing frustum). Additionally, each geometric primitive can be scaled based on the depth of the viewing frustum. Then all potentially visible geometric primitives are transferred to the rasterization stage 660.
[0077] The rasterization stage 660 converts 3D geometric primitives into 2D fragments (e.g., that can be used for display, etc.). The rasterization stage 660 can be configured to set a set of plane equations using the vertices of the geometric primitive, from which various attributes can be interpolated. The rasterization stage 660 can also calculate a coverage mask for multiple pixels, which indicates whether one or more sampling positions of the pixels intercept the geometric primitive. In one embodiment, a z-test can also be performed to determine whether the geometric primitive is occluded by other geometric primitives that have already been rasterized. The rasterization stage 660 generates fragment data (i.e., interpolated vertex attributes associated with specific sampling positions of each covered pixel), which is transferred to the fragment shading stage 670.
[0078] The fragment shading stage 670 processes the fragment data by performing a set of operations (i.e., fragment shader or program) on each of the fragments. The fragment shading stage 670 can generate pixel data (i.e., color values) for the fragments, such as by performing lighting operations or sampling texture maps using the interpolated texture coordinates of the fragments. The fragment shading stage 670 generates pixel data that is sent to the raster operations stage 680.
[0079] In one embodiment, the fragment shading stage 670 can sample texture maps using one or more texture units 345 of the PPU 200. Texture data 603 can be read from the memory 204, and the texture data 603 can be sampled using the texture unit 345 hardware. The texture unit 345 can return the sampled values to the fragment shading stage 670 for processing by the fragment shader.
[0080] The raster operation stage 680 can perform various operations on the pixel data, such as performing an alpha test, a stencil test, and blending the pixel data with other pixel data corresponding to other fragments associated with the pixel. When the raster operation stage 680 has finished processing the pixel data (i.e., the output data 602), the pixel data can be written to a render target, such as a frame buffer, a color buffer, etc.
[0081] It should be appreciated that one or more additional stages can be included in the graphics processing pipeline 600 in addition to or in place of one or more of the above stages. Various implementations of the abstract graphics processing pipeline can implement different stages. Additionally, in some embodiments, one or more of the above stages can be excluded from the graphics processing pipeline (such as the geometry shader stage 640). Other types of graphics processing pipelines are considered to be within the scope of what is contemplated by this disclosure. Additionally, any stage of the graphics processing pipeline 600 can be implemented by one or more dedicated hardware units within a graphics processor (such as the PPU 200). Other stages of the graphics processing pipeline 600 can be implemented by programmable hardware units (such as the SM 340 of the PPU 200).
[0082] The graphics processing pipeline 600 can be implemented via an application executed by a host processor (such as the CPU 550). In one embodiment, a device driver can implement an application programming interface (API) that defines various functions that can be utilized by an application to generate graphics data for display. The device driver is a software program that includes a plurality of instructions for controlling the operation of the PPU 200. The API provides an abstraction to the programmer that allows the programmer to utilize the dedicated graphics hardware (such as the PPU 200) to generate graphics data without requiring the programmer to utilize the specific instruction set of the PPU 200. The application can include API calls that are routed to the device driver of the PPU 200. The device driver interprets the API calls and performs various operations in response to the API calls. In some cases, the device driver can perform operations by executing instructions on the CPU 550. In other cases, the device driver can perform operations by at least partially initiating operations on the PPU 200 by utilizing the input / output interface between the CPU 550 and the PPU 200. In one embodiment, the device driver is configured to utilize the hardware of the PPU 200 to implement the graphics processing pipeline 600.
[0083] Various programs can be executed within the PPU 200 to implement the various stages of the graphics processing pipeline 600. For example, a device driver can start a kernel on the PPU 200 to execute the vertex shading stage 620 on one SM 340 (or multiple SM 340s). The device driver (or an initial kernel executed by the PPU 200) can also start other kernels on the PPU 200 to execute other stages of the graphics processing pipeline 600, such as the geometry shading stage 640 and the fragment shading stage 670. Additionally, some of the stages of the graphics processing pipeline 600 can be implemented on fixed unit hardware, such as a rasterizer or a data assembler implemented within the PPU 200. It should be appreciated that the results from one kernel can be processed by one or more intermediate fixed function hardware units before being processed by subsequent kernels on the SM 340.
[0084] Matrix multiply-accumulate (MMA) operation
[0085] The MMA operation extends the concept of the FMA operation to matrix input operands. In other words, many algorithms are designed around the basic arithmetic operation of multiplying a first input matrix by a second input matrix and adding the result to a third input matrix (i.e., the accumulator matrix). More specifically, the MMA operation can take two input matrices (A and B) and a third accumulator matrix (C in ) to perform the following operation:
[0086] C out = A * B + C in (Equation 1)
[0087] where A is an input matrix of size N×K, B is an input matrix of size K×M, and C is an accumulator matrix of size N×M. The accumulator matrix C is read from a register file, and the result of the MMA operation is accumulated and written over the data of the accumulator matrix C in the register file. In one embodiment, the accumulator matrix C and the result matrix D (C out = D) can be different operands such that the result of the MMA operation is not written on the accumulator matrix C.
[0088] Figure 7 illustrates an MMA operation according to one embodiment. The MMA operation multiplies the input matrix A 710 by the input matrix B 720 and accumulates the result into the accumulator matrix C 730. As Figure 7 shown, the input matrix A is given as an 8×4 matrix, the input matrix B is given as a 4×8 matrix, and the accumulator matrix C is given as an 8×8 matrix. In other words, Figure 7 the MMA operation shown in corresponds to (1) N = 8; (2) M = 8; and (3) K = 4. However, Figure 7Nothing shown herein should be construed as limiting the MMA operations to these dimensions. In fact, the data path of the processor can be designed to operate on matrix operands of any size, as will be shown in more detail below, and matrix operands that are not exactly aligned with the fundamental size of the vector inputs in the dot product operation can be reduced to multiple intermediate operations using the data path.
[0089] Now returning to Figure 7 , each element of the matrix operand can be a value encoded in a particular format. Various formats include but are not limited to single-precision floating-point values (e.g., 32-bit values encoded according to the IEEE 754 standard); half-precision floating-point values (e.g., 16-bit values encoded according to the IEEE 754 standard); signed / unsigned integers (e.g., 32-bit two's complement integers); signed / unsigned short integers (e.g., 16-bit two's complement integers); fixed-point formats; and other formats.
[0090] In one embodiment, the processor can be designed for a 64-bit architecture such that data words are stored in registers having a 64-bit width. Typically, the processor will then implement a data path that operates on values encoded using up to 64-bit formats; however, some data paths can be designed to operate on values encoded using fewer bits. For example, a vector machine can be designed to pack two or four elements, each encoded in 32 bits or 16 bits, respectively, into each 64-bit register. The data path is then configured to execute the same instruction on multiple elements of the parallel input vectors on multiple similar vector units. However, it should be appreciated that vector machines typically perform operations on the elements of the input vectors as completely independent operations. In other words, each of the elements packed into a single 64-bit register is used for only one vector operation and is not shared between different vector units.
[0091] In one embodiment, each element of the input matrix A 710 and each element of the input matrix B 720 can be encoded as a half-precision floating-point value. If each data word is 64 bits wide, then four elements of the input matrix can be packed into each data word. Thus, each register in the register file allocated to store at least a portion of the input matrix A 710 or the input matrix B 720 has the ability to store four half-precision floating-point elements of the corresponding input matrix. This enables the efficient storage of matrix operands to be implemented in a common register file associated with one or more data paths of the processor.
[0092] It should be appreciated that the present invention is not limited to half-precision floating-point data. In some embodiments, each element of the input matrix may be encoded as a full-precision floating-point value. In other embodiments, each element of the input matrix may be encoded as a 16-bit signed integer. In other embodiments, the elements of the input matrix A 710 may be encoded as half-precision floating-point values, while the elements of the input matrix B 720 may be encoded as 32-bit signed integers. In such embodiments, the elements of either input operand may be converted from one format to another in the first stage of the data path such that the formats of each of the input operands may be mixed within a single matrix multiply and accumulate operation. Additionally, in another embodiment, the elements of the input matrix A 710 and the input matrix B 720 may be encoded as half-precision floating-point values, while the elements of the collector matrix C 730 may be encoded as full-precision floating-point values. The data path may even be designed to work with elements of the collector matrix C 730 that have a different precision than the elements of the input matrix A 710 and the input matrix B 720. For example, the accumulator registers in the data path may be extended to store the elements of the collector matrix C 730 as full-precision floating-point values, which add the initial values of the elements of the collector matrix C 730 to the result of a dot product operation performed on half-precision floating-point values, which may be equivalent to the full-precision floating-point values of the partial products if the multiplication is performed in a lossless manner.
[0093] As Figure 7 shown, the matrix has been divided into visually 4×4 element submatrices. In an embodiment, where each element of the input matrix A 710 and the input matrix B 720 is encoded as a half-precision floating-point value (e.g., 16-bit floating point), the 4×4 element submatrices are substantially four 4-element vectors from the matrix. In the case of the input matrix A 710, the matrix is divided into an upper set of vectors and a lower set of vectors. Each vector may correspond to a row of the input matrix A 710, where each row of four elements may be packed into a single 64-bit register. In the case of the input matrix B 720, the matrix is divided into a left set of vectors and a right set of vectors. Each vector may correspond to a column of the input matrix B 720, where each column of four elements may be packed into a single 64-bit register. In the case of the collector matrix C 730, the matrix is divided into four 4×4 element submatrices defined as the upper left quadrant, the upper right quadrant, the lower left quadrant, and the lower right quadrant. As long as the elements are encoded as half-precision floating-point values, each quadrant stores four 4-vector elements from the collector matrix C 730. Each quadrant may correspond to multiple vectors (i.e., a portion of a row or a portion of a column) of the collector matrix C 730. Each quadrant also corresponds to multiple dot product operations performed using the corresponding pairs of vectors from the input matrix.
[0094] For example, as Figure 7 shown, the collector matrix C 0,0The first element is the result of the dot product operation between the first vector <A 0,0 , A 0,1 , A 0,2 , A 0,3 > of the input matrix A710 and the first vector <B 0,0 , B 1,0 , B 2,0 , B 3,0 > of the input matrix B720. The first vector of the input matrix A710 represents the first row of the input matrix A710. The first vector of the input matrix B720 represents the first column of the input matrix B720. Therefore, the dot product between these two vectors is given by:
[0095] C 0,0 = A 0,0 B 0,0 + A 0,1 B 1,0 + A 0,2 B 2,0 + A 0,3 B 3,0 + C 0,0 (Equation 2) where the dot product operation is basically four multiplication operations performed on the corresponding elements of two vectors followed by the execution of four addition operations, and the four addition operations sum the four partial products generated by the multiplication operations with the initial values of the elements of the collector matrix. Then, using different combinations of the vectors of the input matrix, each of the other elements of the collector matrix C730 is calculated in a similar manner. For example, another element (element C 3,2 ) of the collector matrix C730 is generated as the result of the dot product operation between the fourth vector <A 3,0 , A 3,1 , A 3,2 , A 3,3 > of the input matrix A710 and the third vector <B 0,2 , B 1,2 , B 2,2 , B 3,2 > of the input matrix B720. As Figure 7As shown by the MMA operation of, each vector of the input matrix A 710 is consumed by eight dot product operations, and the eight dot product operations are configured to generate the corresponding rows of the elements of the collector matrix C 730. Similarly, each vector of the input matrix B 720 is consumed by eight dot product operations, and the eight dot product operations are configured to generate the corresponding columns of the elements of the collector matrix C 730. Although each of the 64 dot product operations for generating the elements of the collector matrix C 730 is uniquely defined by using different vector pairs from the input matrices, each vector of the first input operand and each vector of the second input operand are consumed by multiple dot product operations and contribute to multiple individual elements of the result matrix.
[0096] It should be appreciated that the above MMA operation can be accelerated by loading the vector sets from the two input matrices into the input of the data path as long as the data path can be configured to consume the vector sets in an efficient manner to simplify the bandwidth between the register file and the input of the data path. For example, in one embodiment, the first two rows of the upper left quadrant of the collector matrix C 730 can be calculated by a data path configured to receive as inputs the first two vectors in the upper vector set of the input matrix A 710 and the first four vectors in the left vector set of the input matrix B 720, as well as the first two vectors (i.e., rows) of the upper left quadrant of the collector matrix C 730. Such a data path would require 8 64-bit words as input: two 64-bit words to store the two vectors of the input matrix A 710, four 64-bit words to store the four vectors of the input matrix B 720, and two 64-bit words to store the two vectors of the collector matrix C 730. Additionally, if the elements of the collector matrix C 730 are encoded as full-precision floating point values (e.g., 32-bit floating point), the size of the input to the data path for the two vectors of the collector matrix C 730 would be doubled to four 64-bit words.
[0097] The data path can then be configured to perform eight dot product operations in parallel in a single pass, a serial multiple pass, or a combination of serial and parallel operations. For example, the data path can be designed to perform one 4-vector dot product operation per pass, which obtains one vector from the input matrix A 710 and one vector from the input matrix B 720, and generates a single element of the collector matrix C 730. Then, over 8 passes, the data path is operated on over 8 passes with different combinations of 6 vectors from the two input matrices to generate 8 different elements of the collector matrix C 730. Optionally, the data path can be designed to perform four 4-vector dot product operations per pass, which obtains one vector from the input matrix A 710 and four vectors from the input matrix B 720, and generates 4 elements of the collector matrix C 730 in parallel. Then, during each pass, the data path is operated on over 2 passes with different vectors from the input matrix A 710 and the same four vectors from the input matrix B 720 to generate 8 elements of the collector matrix C 730. It should be appreciated that the inputs to the data path can be loaded once from the register file before multiple dot product operations are performed by the data path using different combinations of the inputs in each dot product operation. This significantly reduces the bandwidth between the register file and the data path. For example, only 6 vectors in the two input matrices A and B need to be loaded from the register file into the inputs of the data path in order to perform eight dot product operations, whereas using the data path to perform all eight dot product operations individually, where the data path is capable of performing a single dot product operation and only has an input capacity for two vectors, would require 16 vectors to be loaded from the register file into the inputs of the data path because the vectors are reused in multiple dot product operations.
[0098] It should be appreciated that even if the data path is configured to generate dot products of different lengths for each of the vectors in the vector (i.e., the dimension K of the input matrix is not equal to the number of partial products within the data path for a single dot product operation), the data path uses an accumulator register (e.g., collector matrix C 730) such that each vector can be divided into multiple sub-vectors and then loaded into the input of the data path over multiple execution cycles (where after each cycle, the output of the collector matrix C 730 is reloaded into the input of the data path for the next cycle). Thus, the dimension K of the input matrices A 710 and B 720 is not limited to a particular implementation of the dot product operation performed by the data path. For example, if the data path only produces 2-vector dot products (i.e., dot products corresponding to a pair of two-element vectors), then each row of the input matrix A 710 can be divided into a first vector of the first half of the row and a second vector of the second half of the row, and each column of the input matrix B 720 can be divided into a first vector of the upper half of the column and a second vector of the lower half of the column. Then, the elements of the collector matrix C 730 are generated over multiple instruction cycles, where the first half of the vectors of the input matrix A 710 and the upper half of the vectors of the input matrix B 720 are loaded into the input of the data path during the first instruction cycle, and the second half of the vectors of the input matrix A 710 and the lower half of the vectors of the input matrix B 720, along with the intermediate results stored in the collector matrix C 730 during the first instruction cycle, are loaded into the input of the data path during the second instruction cycle. By dividing each of the vectors in the input matrix into multiple parts, each part having a number of elements (equal to the size of the dot product operation implemented by the data path), the MMA operation can be simplified in this way for input matrices of any size dimension K. Even if the dimension K is not divisible by the size of the dot product operation, the vectors can be padded with zeros to obtain the correct result.
[0099] Figure 8 is a conceptual diagram of a dot product operation according to an embodiment. The dot product operation basically adds multiple partial products. The dot product operation can specify three operands, vector A, vector B, and scalar collector C. Vectors A and B have the same length (i.e., the number of elements). As Figure 8 shown, the lengths of vectors A and B are given as 2; however, it should be appreciated that the dot product operation can have any length greater than or equal to 2.
[0100] The dot product operation multiplies pairs of elements from the input vectors A and B. As Figure 8As shown, in multiplier 822, the first element A0812 of input vector A is multiplied by the corresponding element B0814 of input vector B to generate partial product A0B0826. In multiplier 824, the second element A1816 of input vector A is multiplied by the corresponding element B1818 of input vector B to generate partial product A1B1828. Then, three-element adder 830 is used to sum partial product A0B0826, partial product A1B1828, and scalar collector value C in 820 to generate result value C out 832. Result value C out 832 can be stored in the register for scalar collector value C in 820 and can be reused to accumulate multiple dot product operations for longer vectors.
[0101] In addition, the dot product operation can be extended by adding additional multipliers 822, 824, etc. in parallel to compute additional partial products, and then summing the additional partial products with either a larger element adder or a tree of smaller adders that generate intermediate sums, and then summing again by an additional multi-element adder.
[0102] Although the dot product operation can be implemented in a traditional FMA data path where each partial product is computed and accumulated into an accumulation register during one pass through the data path, it is more efficient to compute multiple partial products of the dot product operation in parallel and sum the results in a single multi-stage pipeline. In addition, although multiple cores can be used simultaneously in SIMD / SIMT machines to compute partial products in parallel, an additional step of summing all partial products is still required, which is not trivial to accomplish efficiently in such machines.
[0103] Figure 9 A portion of a processor 900 including a data path 930 configured to implement matrix operations according to one embodiment is shown. Processor 900 can refer to a central processing unit (CPU), a graphics processing unit (GPU), or other parallel processing unit, a reduced instruction set computer (RISC)-type processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), etc. Nothing in this disclosure should be construed as limiting processor 900 to a parallel processing unit such as PPU 200.
[0104] As Figure 9As shown, the processor 900 includes a multi-bank register file implemented as a plurality of register banks 910. Each register bank 910 can store multiple data words in multiple registers. Each register bank 910 can have independent and different read and write ports, such that one register in the register bank 910 can be read and another register can be written in any given clock cycle. Thus, one data word can be simultaneously read from each register bank 910 and loaded into the operand collector 920 during a single clock cycle. The register file is configured to store the operands specified in the instructions for MMA operations. In one embodiment, each operand specified in the instructions is a matrix having a plurality of elements in a two-dimensional array of rows and columns, and each register can store one or more elements of a particular operand. Of course, in one embodiment, the register file can include only a single bank, meaning that only one register can be read from the register file and loaded into the operand collector 920 each clock cycle.
[0105] The processor 900 also includes a plurality of operand collectors coupled to the inputs of one or more data paths. In one embodiment, the operand collector 920 includes a plurality of asynchronous flip-flops, which enable data to be loaded into the operand collector 920 during any particular clock cycle and then read from the operand collector 920 in any subsequent clock cycle. In other words, based on the signals at the inputs of the flip-flops, the flip-flops are not set or reset every clock cycle. Instead, the control logic determines when the flip-flops are set or reset and when the data stored in the flip-flops is transferred to the outputs of the flip-flops based on the input signals. This enables the plurality of operand collectors 920 to load operands from the register file over multiple clock cycles and then provide the plurality of operands to the inputs of the data path in parallel during a single clock cycle. It should be appreciated that the operand collector 920 can be implemented in a variety of different ways, including various different types of latches and / or flip-flops, and various embodiments can use different underlying technologies to implement the operand collector 920. However, the function of the operand collector 920 is to temporarily store the operands required to perform operations on the data path, where the operands can be loaded from the register file 910 over one or more clock cycles, depending on which register bank 910 the operands are stored in and how many read ports are available in those register banks 910.
[0106] A crossbar switch 915 or other type of switchable interconnect can be coupled to the read ports of the register bank 910 and the inputs of the operand collectors. The crossbar switch 915 can be configured to route signals from the read ports associated with any register bank 910 to a particular operand collector 920. For example, the read port of register bank 1 910(1) can include 64 interconnects carrying signals corresponding to 64 bits contained in a single register of the register file. These 64 interconnects can be connected to one of a plurality of different operand collectors 920, each of which includes 64 flip-flops for storing the 64 bits encoded by the signals transmitted via the read port. If the data path requires three operand collectors 920 coupled to the input of the data path, and each operand collector 920 includes 64 flip-flops to store the 64 bits of the corresponding operand of the data path, then the crossbar switch 915 can be configured to route the 64 signals on the 64 interconnects of the read port to any of the three operand collectors 920.
[0107] The operand collectors 920 can be coupled to the inputs of one or more data paths. As Figure 9 shown, the operand collectors 920 can be coupled to a half-precision matrix multiply-accumulate (HMMA) data path 930 and a double-precision (64-bit) floating-point (FP64) data path 940. The FP64 data path 940 can be a conventional double-precision floating-point FMA data path that enables addition, subtraction, multiplication, division, and other operations to be performed on double-precision floating-point operands. In one embodiment, the FP64 data path 940 can include logic for performing an FMA operation on three scalar double-precision floating-point operands (e.g., A, B, and C), as is well known in the art.
[0108] The output of the FP64 data path 940 is coupled to the result queue 950. The result queue 950 stores the results generated by the FP64 data path 940. In one embodiment, the result queue 950 includes a plurality of flip-flops for storing a plurality of bits of the results generated by the FP64 data path 940. For example, the result queue 950 may include 64 flip-flops for storing the double-precision floating-point results of the FP64 data path 940. The result queue 950 enables the temporary storage of the results while waiting for the availability of the write port to write the value back to the register file. It should be appreciated that the FP64 data path 940 may be included in each of a plurality of similar cores of the shared multi-bank register file 910. During a particular clock cycle, only one core can write a value back to each register bank. Thus, if two or more cores generate results during a given clock cycle and both results need to be written back to the same register bank 910, one result can be written to the register bank 910 during the first clock cycle and the other result can be written to the register bank 910 during the second clock cycle.
[0109] It should be appreciated that the result queue 950 may be attached to an accumulator register included within the data path that does not need to be written back to the register file between the execution of multiple instructions. For example, an FMA instruction may include operands A, B, and C during a first instruction, and then only operands A and B during one or more subsequent instructions, chaining multiple instructions together using an internal accumulator register and reusing the accumulated value C as the third operand for each subsequent instruction. In some embodiments, if the results generated by the FP64 data path 940 are always written immediately to the register file as soon as they become available, the result queue 950 may be omitted. However, such an architecture requires more advanced control of the memory allocation of the multi-bank register file to avoid any conflicts with the write ports, as well as knowledge of the pipeline lengths of two or more cores sharing the register file, to appropriately schedule which cores will need to access a given write port during a particular clock cycle. In a processor with a large number of cores, it may be easier to write values back to the register file using the result queue 950 as needed, causing any core to stall from completing subsequent instructions until the result has been written back to the register file.
[0110] In one embodiment, the HMMA data path 930 shares the same available operand collector 920 with the FP64 data path 940. The HMMA data path 930 and the FP64 data path 940 may be included in a common core of the processor 900, which includes multiple cores, each core including an FP64 data path 940 and an HMMA data path 930 and possibly also an integer arithmetic logic unit (ALU). In one embodiment, the HMMA data path 930 is configured to perform matrix multiply-accumulate (MMA) operations. Instructions for the MMA operations specify multiple matrix operands, which are configured to perform operations equivalent to the function specified by Equation 1 above.
[0111] In one embodiment, the multiple operand collectors 920 include storage for at least two vectors of a first operand (i.e., input matrix A 710) specified in the instruction and at least two vectors of a second operand (i.e., input matrix B 720) specified in the instruction. Each of the at least two vectors has at least two elements in a row or column of the matrix operand. For example, in one embodiment, the HMMA data path 930 is configured to receive two vectors from the first operand and four vectors from the second operand as inputs to the data path. Thus, the number of operand collectors 920 should be sufficient to store at least six vectors of two input matrix operands (e.g., at least six 64-bit operand collectors). Depending on the design of the HMMA data path 930, other embodiments may require more or fewer operand collectors 920.
[0112] In one embodiment, the HMMA data path 930 is further configured to receive at least two vectors of a third operand (i.e., the collector matrix C 730) specified in the instruction. The collector matrix C 730 is added to the result of multiplying the first and second operands specified in the instruction. The number of combined elements in the vectors from the third operand must match the product of the number of vectors of the first operand and the number of vectors of the second operand. For example, if the plurality of operand collectors 920 store two vectors (e.g., rows) of the input matrix A 710 and four vectors (e.g., columns) of the input matrix B 720, then the number of elements in at least two vectors of the collector matrix C 730 must equal eight. Additionally, the indices of the elements of the third operand must match the indices of the vectors of the first and second operands. For example, if the two vectors of the first operand correspond to the first and second rows of the input matrix A 710, and the four vectors of the second operand correspond to the first through fourth rows of the input matrix B 720, then the indices of the elements of the third operand must match the <row, column> index vectors associated with the dot product of any vector of the input matrix A 710 and any vector of the input matrix B 720, where the indices of the elements of the third operand are two-dimensional.
[0113] Similarly, the HMMA data path 930 generates a plurality of elements of the result matrix at the output of the HMMA data path 930. Each element of the plurality of elements of the result matrix is generated by computing at least one dot product of corresponding vector pairs selected from the matrix operands. The dot product operation may include the step of accumulating a plurality of partial products into the result queue 950. Each partial product of the plurality of partial products is generated by multiplying each element of the first vector by the corresponding element of the second vector. An example of the dot product operation is given in Equation 2 above. It should be appreciated that in one embodiment, a plurality of partial products are computed in parallel in the HMMA data path 930, and an adder tree internal to the HMMA data path 930 is used to accumulate the plurality of partial products before output to the result queue 950. In yet another embodiment, the partial products are computed serially in multiple passes within the HMMA data path 930, and each partial product is accumulated into an accumulator register internal to the HMMA data path 930. When all of the partial products from the collector matrix C 730, as well as the addend value, have been accumulated into the internal accumulator register in multiple passes, the final result is output to the result queue 950.
[0114] In one embodiment, the processor 900 is implemented as the PPU 200. In such an embodiment, each core 450 in the SM 340 includes an HMMA data path 930, an FP64 data path 940, and optionally includes an integer ALU. The register file 420 may implement one or more memory banks 910. The crossbar 915 and the operand collector 920 may be implemented between the register file 420 and one or more cores 450. In addition, the result queue 950 may be implemented between one or more cores 450 and the interconnect network 480, which enables the results stored in the result queue 950 to be written back to the register file 420. Thus, the processor 900 is the PPU 200 including multiple SM 340s, each of the multiple SM 340s includes a register file 420 and multiple cores 450, and each of the multiple cores 450 includes an instance of the HMMA data path 930.
[0115] The PPU 200 implements a SIMT architecture, which enables parallel execution of multiple threads on multiple cores 450 in multiple SM 340s. In one embodiment, the MMA operation is configured to execute multiple threads in parallel on multiple cores 450. Each thread is configured to generate a portion of the elements in the result matrix (e.g., the collector matrix C 730) on a specific core 450 using a different combination of vectors of the operands specified in the instruction for the MMA operation.
[0116] For example, as Figure 7 shown, the MMA operation can be executed on the 8×4 input matrix A 710 and the 4×8 input matrix B 720 simultaneously on 8 threads. The first thread is assigned to the first two vectors of the input matrix A 710 (e.g., <A 0,0 ,A 0,1 ,A 0,2 ,A 0,3 > and <A 1,0 ,A 1,1 ,A 1,2 ,A 1,3 >) and the first four vectors of the input matrix B 720 (e.g., <B 0,0 ,B 1,0 ,B 2,0 ,B 3,0 >, <B 0,1 ,B 1,1 ,B 2,1 ,B 3,1 >, <B 0,2 ,B 1,2 ,B 2,2 ,B 3,2 > and <B 0,3 ,B 1,3 ,B2,3 , B 3,3 >). The first thread generates eight elements included in two vectors of the result matrix (e.g., <C 0,0 , C 0,1 , C 0,2 , C 0,3 > and <C 1,0 , C 1,1 , C 1,2 , C 1,3 >). Similarly, the second thread is assigned to the first two vectors of input matrix A710 (e.g., <A 0,0 , A 0,1 , A 0,2 , A 0,3 > and <A 1,0 , A 1,1 , A 1,2 , A 1,3 >) and the next four vectors of input matrix B720 (e.g., <B 0,4 , B 1,4 , B 2,4 , B 3,4 , <B 0,5 , B 1,5 , B 2,5 , B 3,5 , <B 0,6 , B 1,6 , B 2,6 , B 3,6 , and <B 0,7 , B 1,7 , B 2,7 , B 3,7 >). The second thread generates eight elements included in two different vectors of the result matrix (e.g., <C 0,4 , C 0,5 , C 0,6 , C 0,7 > and <C 1,4 , C 1,5 , C 1,6 , C 1,7 >). The third thread is assigned to the next two vectors of input matrix A710 (e.g., <a 2,0 , a 2,1 , A 2,2 , A 2,3 > and <A 3,0 , A 3,1 , A 3,2 , A 3,3 >) and the first four vectors of input matrix B720 (e.g., <B 0,0 , B 1,0 , B 2,0 , B 3,0,<B 0,1 ,B 1,1 ,B 2,1 ,B 3,1 ,<B 0,2 ,B 1,2 ,B 2,2 ,B 3,2 > and <B 0,3 ,B 1,3 ,B 2,3 ,B 3,3 >). The third thread generates eight elements that are included in two vectors of the result matrix (e.g., <C 2,0 ,C 2,1 ,C 2,2 ,C 2,3 > and <C 3,0 ,C 3,1 ,C 3,2 ,C 3,3 >). The other five threads perform similar operations in combination with additional vectors from input matrix A710 and input matrix B 720.
[0117] It should be appreciated that each thread is assigned to a core 450, and the vectors assigned to that thread are loaded into the operand collector 920 of the core 450, and then the elements of the result matrix are generated by performing MMA operations on the HMMA data path within the core 450. In one embodiment, each core is coupled to a dedicated set of operand collectors 920 that are coupled only to that core 450. In another embodiment, multiple cores 450 share the operand collector 920. For example, two cores 450 having two HMMA data paths 930 can share a set of operand collectors 920, where the common vectors assigned to two threads scheduled on the two cores 450 are shared by the two cores 450. In this way, the common vectors assigned to two or more threads are not loaded into two separate sets of operand collectors 920. For example, the first two vectors of input matrix A 710 are assigned to both of the above-mentioned first two threads, while different sets of vectors of input matrix B 720 are assigned. Therefore, the operand collectors 920 for storing the vectors of input matrix A710 can be shared between the two cores 450 by coupling these operand collectors 920 to the inputs of the two HMMA data paths 930.
[0118] It should be appreciated that any number of threads can be combined to increase the size of the MMA operation. In other words, by adding more threads to process additional computations in parallel, the dimensions M, N, and K of the MMA operation can be increased without increasing the execution time. Optionally, increasing the size of the MMA operation can also be accomplished on a fixed number of cores 450 by executing multiple instructions on each core 450 over multiple instruction cycles. For example, a first thread can be executed on a particular core 450 during a first instruction cycle, and a second thread can be executed on the particular core 450 during a second instruction cycle. Since the vectors of the input matrix A 710 are shared between the first and second threads and thus the vectors of the input matrix A 710 do not need to be reloaded from the register bank 910 into the operand collector 920 between two instruction executions over two instruction cycles, it may be beneficial to execute the MMA operation over multiple instruction cycles.
[0119] Figure 10 A conventional double-precision floating-point FMA data path 1000 according to one embodiment is shown. The conventional double-precision floating-point FMA data path 1000 shows one possible implementation of the FP64 data path 940 of the processor 900. The data path 1000 implements an FMA operation that takes three operands (A, B, and C) as inputs, multiplies operand A by operand B, and adds the product to operand C. Each of the three operands is a double-precision floating-point value encoded with 64 bits: 1 sign bit, 11 exponent bits, and 52 mantissa bits.
[0120] As Figure 10 shown, the data path 1000 includes a multiplier 1010 that multiplies the mantissa bits from the A operand 1002 by the mantissa bits from the B operand 1004. In one embodiment, the multiplier 1010 can be, for example, a Wallace tree that multiplies each bit of one mantissa by each bit of the other mantissa and combines the partial products with an adder tree in multiple reduction layers to generate two n-bit integers that are added to obtain the binary result of the multiplication. If the multiplier is designed for 64-bit floating-point numbers, thus multiplying two 52-bit mantissas plus a hidden bit (for a normalized value), then the multiplier 1010 can be a 53×53 Wallace tree that produces two 106-bit values at the output, which are added in a 3:2 carry sum adder (CSA) 1040.
[0121] In parallel, the exponent bits from operand A 1012 are added to the exponent bits from operand B 1014 in adder 1020. Adder 1020 can be a full adder instead of a CSA adder because the exponent bits are only 11 bits wide, and the full addition can be propagated through adder 1020 in a similar time as the result of multiplier 1010 is propagated through the reduction layers of the Wallace tree to generate two integers. The result produced by adder 1020 gives the exponent associated with the result of the product of the mantissa bits. Then, based on the exponent bits of operand C 1016, the exponent associated with the result is used to shift the mantissa bits from operand C 1006. It should be appreciated that the exponents must be the same when adding the mantissa bits, and incrementing or decrementing the exponent of a floating-point value is equivalent to shifting the mantissa left or right. Then, the shifted mantissa bits of operand C 1006 are added to the two integers generated by multiplier 1010 in 3:2 CSA 1040. 3:2 CSA 1040 generates a carry value and a sum value, which represent the addition result, where each bit of the sum value represents the result of adding three corresponding bits (one bit from each of the three inputs), and each bit of the carry value represents a carry bit, which indicates whether the addition of these three corresponding bits results in a carry (i.e., the bit that needs to be added to the next most significant bit in the sum value). 3:2 CSA 1040 enables all carry bits and sum bits to be calculated immediately without having to propagate the carry bits to each subsequent three-bit addition operation.
[0122] Then, the carry value and the sum value from 3:2 CSA 1040 are added in finish adder 1050. The result produced by finish adder 1050 represents the sum of the three mantissa values from the three operands. However, this sum is not normalized while the floating-point value is normalized. As a result, the result produced by finish adder 1050 will be shifted by normalization logic 1060 by a certain number of bits such that the most significant bit in the result is 1, and the exponent produced by adder 1020 will be incremented or decremented by normalization logic 1060 accordingly based on the number of bits the result is shifted. Finally, the normalized result is rounded by rounding logic 1070. The result produced by finish adder 1050 is much larger than 52 bits. Since the result cannot be losslessly encoded in the 52 mantissa bits of a double-precision floating-point value, the result is rounded so that the mantissa bits of the result are only 52 bits wide.
[0123] It should be appreciated that Figure 10 the illustration of the sign logic is omitted, but the operations of the sign logic and the conventional floating-point data path are well understood by those skilled in the art and should be considered to be within the scope of data path 1000. The sign bit, the normalized exponent bit, and the rounded mantissa bit are output by data path 1000 and stored as a double-precision floating-point result in operand C 1008.
[0124] Figure 11 Illustrates the HMMA data path 1100 according to one embodiment. The HMMA data path 1100 includes a pair of half-precision floating-point FMA units 1110. Similar to the data path 1000, each of the units 1110 implements an FMA operation: taking three operands (A, B, and C) as inputs, multiplying operand A by operand B and adding the product to operand C. However, different from the data path 1000, each of the three operands is a half-precision floating-point value encoded with 16 bits: 1 sign bit, 5 exponent bits, and 10 mantissa bits. Except that the components of the unit 1110 are significantly smaller than the similar components of the data path 1000, since the number of bits in each operand is reduced from 64 bits to 16 bits, the unit 1110 is similar to the data path 1000 in implementation. Thus, the multiplier of the data path 1100 can be implemented as an 11×11 Wallace tree instead of a 53×53 Wallace tree. Similarly, the sizes of the 3:2 CSA adder, the carry adder, the normalization logic, and the rounding logic are reduced by approximately 1 / 4. Otherwise, the functional description of the data path 1000 applies equally well to the half-precision floating-point FMA unit 1100, only on the operands represented with less bit precision.
[0125] It should be appreciated that each unit 1110 is used to multiply two half-precision floating-point values from two input operands and add the product to the addend from the third input operand. Thus, each unit 1110 can be used in parallel to calculate partial products of a dot product operation. In one embodiment, the first unit 1110(0) is provided with one element from each of two input vectors and and the second unit 1110(1) is provided with another element from each of the two input vectors and where each input vector includes two elements. For example, the first unit 1110(0) is provided with elements A0 and B0 of the input vectors and respectively. The size of the dot product operation corresponds to the number of units 1110 implemented in parallel. However, summing the partial products generated by each of the units 1110 requires additional combinational logic.
[0126] It should be appreciated that if the HMMA data path 1100 is implemented as a vector machine, the combinational logic can be ignored, and each unit 1110 can perform FMA operations on scalar half-precision floating-point values to generate two FMA results at the respective outputs of each unit 1110. In fact, in some embodiments, the HMMA data path 1100 can be configured to do exactly this. However, additional combinational logic is needed to implement the dot product operation and generate a single result using the multipliers in the two units. Thus, the HMMA data path 1100 can be configured into two operation modes: a first mode and a second mode, where in the first mode each unit 1110 performs FMA operations on vector inputs in parallel, and where in the second mode each unit 1110 generates partial products that are passed to the combinational logic. The combinational logic then adds the partial products to the addend from a third input operand.
[0127] In one embodiment, the combinational logic includes exponent comparison logic 1120, product alignment logic 1130, a carry adder tree including 4:2 CSA 1142 and 3:2 CSA 1144, a completion adder 1150, normalization logic 1160, and rounding logic 1170. The product alignment logic 1130 receives two integers output by the multiplier for each unit 1110 from each unit 1110. The product alignment logic 1130 is controlled by the exponent comparison logic 1120, which receives the exponent associated with the partial product. In one embodiment, each of the units 1110 includes logic, such as an adder 1020, which adds the exponent bits associated with two input operands (e.g., A i of unit i, B i ). Then, the output of the logic equivalent to the adder 1020 is routed from the unit 1110 to the exponent comparison logic 1120. Then, the exponent comparison logic 1120 compares the exponent values associated with the partial products generated by the multipliers in each unit 1110, and uses the difference in exponents to generate a control signal that causes the product alignment logic 1130 to shift one of the partial products generated by the multipliers in each unit 1110. Also, the partial products represent the mantissa of the floating-point value, and thus, the bits of the partial products must first be aligned so that the exponents match before performing the addition operation.
[0128] The shifted partial products are then passed to a 4:2 CSA 1142, which adds four integer values and generates a carry value and a sum value. The output of the 4:2 CSA is passed as two inputs to a 3:2 CSA 1144, which adds the carry value and the sum value to an addend from a third operand C. It should be appreciated that the addend can be in a half-precision floating-point format or a single-precision floating-point format, which is encoded in 32 bits: 1 sign bit, 8 exponent bits, and 23 mantissa bits. Remember that the result of multiplying two 11-bit values (10 mantissa bits plus a leading hidden bit) is a 22-bit value. Thus, even though the partial products are generated based on half-precision floating-point values, the width of the partial products is almost the same as the width of the mantissa of the single-precision floating-point addend from the third operand. Of course, the addend can also be a half-precision floating-point value similar to the elements of the input vectors and .
[0129] The result output by the 3:2 CSA 1144 is passed to a finish adder 1150, which is similar to the finish adder 1050 except for having a smaller width. The result then generated by the finish adder 1150 is passed to a normalization logic 1160 and a rounding logic 1170 to shift and truncate the result. The normalization logic 1160 receives the value of the common exponent for the two partial products after the shift, and shifts the result by incrementing or decrementing the exponent value and shifting the bits of the result left or right until the MSB of the result is 1. Then the rounding logic 1170 truncates the result to fit the width of the mantissa bits of at least one floating-point format. The sign bit, the normalized exponent bit, and the rounded mantissa bits are output by the data path 1100 and stored as a half-precision floating-point value or a single-precision floating-point value in the C operand 1108.
[0130] Returning to the top of the data path 1100, it is evident that the selection logic 1105 is coupled to the inputs of the three operands of each unit 1110. As described above, the two-element vector and the two-element vector plus the scalar operand C can be used to perform a dot product operation using the data path 1100. Although a data path that can be configured to perform a dot product operation is generally more useful than a data path that can only be configured to perform an FMA operation, additional functionality is added by including the selection logic 1105, which makes the MMA operation more efficient when the data path 1100 is coupled to an additional operand collector 920.
[0131] For example, the operand collector 920 coupled to the data path 1100 can include multiple operand collectors 920 sufficient to store at least two input vectors associated with the input matrix A 710 and at least two input vectors associated with the input matrix B 720 plus one or more vectors associated with multiple elements of the collector operand C 730. Selection logic 1105 is then used to select elements from different vectors stored in the operand collector 920 to perform multiple dot product operations in multiple passes of the data path 1100, all of the multiple passes of the data path 1100 being associated with a single instruction cycle of the data path 1100. The selection logic 1105 may include multiple multiplexers and control logic for switching the multiplexers between two or more inputs of each multiplexer.
[0132] For example, during the first pass, a first input vector is selected and the first input vector Each input vector has two elements, and a first element / summand from the collector matrix C is selected to generate a first dot product result. During the second pass, the first input vector is selected and a second input vector and a second element / summand from the collector matrix C to generate a second dot product result. During the third pass, the second input vector is selected and the first input vector and a third element / summand from the collector matrix C to generate a third dot product result. Finally, during the fourth pass, the second input vector is selected and the second input vector and a fourth element / summand from the collector matrix C to generate a fourth dot product result. These results may be stored in a result queue 950 having a width of 64 or 128 bits, depending on whether the dot product results stored in the collector matrix C are encoded as half-precision floating-point values or single-precision floating-point values.
[0133] In another embodiment, the 4:2 CSA 1142 and the 3:2 CSA 1144 may be combined into a 5:2 CSA. Although the actual difference between the two embodiments is minimal, there is a slight difference in the order of how the five parameters are added because the 4:2 CSA is typically implemented as a tree of 3:2 CSAs, and the 5:2 CSA is also typically implemented as a tree of 3:2 CSAs.
[0134] In yet another embodiment, the product alignment logic 1130 can be configured to truncate partial products when shifting the partial products, thereby reducing the size of the CSA configured to sum the aligned partial products. To shift the partial products without truncation, the product alignment logic 1130 would need to output an additional width of the partial products to the 4:2 CSA 1142. To avoid an increase in the width of the partial products, the product alignment logic 1130 can shift the partial products at a wider bit width and then truncate to the MSB before transferring the partial products to the 4:2 CSA 1142. This will result in a reduction in the size of the required CSA for summing the partial products and addends.
[0135] In yet another embodiment, the data path 1100 can be scaled to generate dot products for vectors of more than two elements. Generally, for a pair of p-element vectors, the data path 1100 can include p cells 1110 for computing p partial products and additional combinational logic for combining all the partial products. For example, Figure 11 the logic shown in can be doubled to generate two parts of the dot product, and then an additional layer of simplified combinational logic can be included to combine the sum of the two partial products with the sum of two other partial products.
[0136] It should be appreciated that the addend for the dot product operation of two input vectors can be provided as an input to only one of the cells 1110 in the data path 1100. All other cells 1110 of the data path 1100 should receive a zero constant value as the addend operand for the cell 1110 such that the addend is only added to the dot product result once.
[0137] It should be appreciated that when the data path 1100 is configured to use additional combinational logic to generate the dot product result, the CSA, carry adder, normalization logic, and rounding logic of each cell 1110 are not utilized. However, when the data path 1100 is configured as a vector machine to generate a vector of FMA results, this logic will be used. It should be appreciated that when the data path 1100 is configured to generate the dot product result, the 4:2 CSA adder 1142, 3:2 CSA adder 1144, carry adder 1150, normalization logic 1160, and rounding logic 1170 are very similar compared to the unused logic in each cell, although the precision is different. If possible, it is useful to utilize a portion of the logic within the cell 1110 to perform the same operations as the additional combinational logic.
[0138] Figure 12Shows an HMMA data path 1200 according to another embodiment. The HMMA data path 1200 includes a "small" half-precision floating-point FMA unit 1210 and a "large" half-precision floating-point FMA unit 1220. The small unit 1210 is similar to each of the units 1110 and implements an FMA operation: taking three operands (A, B, and C) as inputs, multiplying operand A by operand B and adding the product to operand C. The large unit 1220 is similar to the small unit 1210 in that the large unit 1220 implements an FMA operation: taking three operands (A, B, and C) as inputs, multiplying operand A by operand B and adding the product to operand C. However, the large unit 1220 includes slightly different logic internally to implement both the half-precision floating-point FMA operation and the combinational logic to implement the dot product operation in combination with the small unit 1210.
[0139] As Figure 12 shown, the partial products generated by the small unit 1210 are output to the first partial product resolver 1232. In one embodiment, the partial product resolver 1232 is a carry adder that combines two integers generated by the multiplier into a final value representing the product of the two mantissas of the first partial product. Similarly, the partial products generated by the large unit 1220 are output to a second partial product resolver 1234 similar to the first partial product resolver 1232. The output of the first partial product resolver 1232 and the first of the two integers of the second partial product generated by the multiplier in the large unit 1220 are coupled to a first switch, and the output of the second partial product resolver 1234 and the second of the two integers of the second partial product generated by the multiplier in the large unit 1220 are coupled to a second switch. The first switch and the second switch control whether the large unit 1220 is configured in a first mode to generate the result of a scalar FMA operation or the large unit 1220 is configured in a second mode to generate the result of a vector dot product operation.
[0140] The outputs of the first switch and the second switch are coupled to product alignment logic 1240, which is configured to shift the passed partial products as inputs via the first switch and the second switch when the large unit 1220 is configured in the second mode. The product alignment logic 1240 is controlled by exponent comparison logic 1245, which operates similar to the exponent comparison logic 1120. If the large unit 1220 is configured in the first mode, the product alignment logic 1240 does not shift either of the two integers passed to the product alignment logic 1240 via the first switch and the second switch. When the small unit 1210 and the large unit 1220 are generating scalar FMA results as a vector machine respectively, no shift is performed because the exponents associated with the two partial products are not relevant.
[0141] The aligned partial products are then passed to the 3:2 CSA 1250, which adds two partial products with the addend from the third input operand. The 3:2 CSA 1250 in the large cell 1220 can be significantly wider (i.e., have greater precision) than the corresponding CSA in the small cell 1210. This is necessary for handling the additional bits of precision required for the lossless operations of the partial product parsers 1232, 1234, and the product alignment logic 1240.
[0142] In one embodiment, the data path 1200 further includes selection logic 1205 for selecting among multiple combinations of at least two vectors from the input matrix A 710, at least two vectors from the input matrix B 720, and different elements / addends from the collector matrix C 730, so as to generate multiple dot product results in the result queue 950.
[0143] Figure 13 An HMMA data path 1300 according to yet another embodiment is shown. The HMMA data path 1300 includes: four multipliers 1310 for generating partial products of two four-element vectors and ; four negation logic 1320 blocks for combining the sign bits of the operands; five shift logic 1330 blocks for shifting the partial products and addends to align all values based on the exponents of the partial products; a CSA tree including multiple reduced layers of 3:2 CSA 1342 and 4:2 CSA 1344; a final adder 1350; normalization logic 1360; and rounding logic 1370. The data path 1300 includes multiple pipeline stages: a first pipeline stage 1301, which includes conversion / encoding logic 1315; a second pipeline stage 1302, which includes multipliers 1310; a third pipeline stage 1303, which includes negation logic 1320, shift logic 1330, and the CSA tree; a fourth pipeline stage 1304, which includes the final adder; a fifth pipeline stage 1305, which includes normalization logic 1360; and a sixth pipeline stage 1306, which includes rounding logic 1370.
[0144] In the first pipeline stage 1301, the conversion / encoding logic 1315 receives the elements of two input vectors and the elements / addends of the collector matrix C 730, and performs one or more preprocessing operations on these elements. The preprocessing may involve converting the elements from one format to a half-precision floating-point value format. For example, the input vectors may be provided in 16-bit floating-point, 8-bit signed / unsigned integer, 16-bit signed / unsigned integer, 32-bit fixed-point format, etc. and The conversion / encoding logic 1315 is configured to convert all input values into a half-precision floating-point value format to be compatible with the rest of the data path 1300.
[0145] In one embodiment, the conversion / encoding logic 1315 may also include a modified Booth encoder. The modified Booth encoder generates selector signals for every three bits of the multiplicand (e.g., bits of the elements of the vector). The selector signals are then passed to the multiplier 1310, which is designed to implement the modified Booth algorithm to generate partial products. The modified Booth algorithm may accelerate the multiplier 1310 by reducing the number of reduction layers (adders) in the multiplier 1310. It should be appreciated that in some embodiments, the data paths 1100 and 1200 may also be modified to incorporate the conversion / encoding logic 1315 and the multiplier designed to implement the modified Booth algorithm.
[0146] In the second pipeline stage 1302, each of the multipliers 1310 receives a corresponding pair of corresponding elements from two input vectors and For example, the first multiplier 1310 receives elements A0 and B0; the second multiplier 1310 receives elements A1 and B1; the third multiplier 1310 receives elements A2 and B2; and the fourth multiplier 1310 receives elements A3 and B3. Each of the multipliers 1310 generates two integers that represent the partial products formed by multiplying the elements input to that multiplier 1310.
[0147] In the third pipeline stage 1303, the negation logic block 1320 combines the sign bits for the elements input to the corresponding multiplier 1310 and, if the combined sign bit is negative, takes the partial product's complement via a two's complement operation applied to the pair of integers. For example, the negation logic block 1320 may perform an exclusive OR (XOR) on the sign bits from the two elements input to the corresponding multiplier 1310. If the partial product is negative, the result of the XOR operation is 1. If the partial product is positive, the result of the XOR operation is 0. If the result of the XOR operation is 1, the two integers from the multiplier 1310 are complemented by performing a two's complement operation on each value (i.e., toggling the state of each bit in the value and then adding 1). It should be appreciated that in some embodiments, the negation logic block 1320 may be implemented in a similar manner in the data paths 1100 and 1200 to handle the sign bits of various operands.
[0148] The shift logic 1330 block shifts the partial products based on the exponents associated with all four partial products. Although not explicitly shown, the exponents associated with each partial product are calculated using an adder (such as adder 1020) by summing the exponent bits included in the elements associated with the respective multiplier 1310. The maximum exponent associated with all four partial products and the addend is provided to each of the shift logic 1330 blocks. Each shift logic 1330 block then determines how many bits to shift the partial product corresponding to that shift logic 1330 block in order to align the partial product with the maximum exponent of all partial products. One of the shift logic 1330 blocks also shifts the addend mantissa. It should be appreciated that the maximum possible shift distance in bits will increase the required bit width of the partial product integers passed to the CSA tree to avoid loss of precision. In one embodiment, the shift logic 1330 blocks are configured to truncate the aligned partial products in order to reduce the precision of the adders in the CSA tree.
[0149] The CSA tree includes multiple reduction levels where three or four inputs are added together to produce two outputs, a carry value and a sum value. As Figure 13 shown, a 4-element dot product operation requires three reduction levels, including three 3:2 CSAs 1342 at the first reduction level; two 3:2 CSAs 1342 at the second reduction level; and one 4:2 CSA 1344 at the third reduction level. The outputs of the 4:2 CSA 1344 at the third reduction level generate the carry value and the sum value of the dot product of two vectors to be added to the addend. In the fourth pipeline stage 1304, the carry value and the sum value are transferred to the finish adder 1350, which sums the carry value and the sum value to produce the mantissa value of the dot product. In the fifth pipeline stage 1305, the result is transferred to the normalization logic 1360, which normalizes the mantissa value of the dot product and adjusts the maximum exponent value of the dot product based on this alignment. In the sixth pipeline stage 1306, the normalized mantissa value and exponent value are transferred to the rounding logic 1370, which rounds the result to the width of the format of the elements of the collector matrix C 730 (e.g., half precision or single precision).
[0150] Although not explicitly shown, selection logic similar to selection logics 1105 and 1205 can be coupled to the transform / encoding logic 1315 such that multiple vectors stored in the operand collector 920 can be used to generate the dot product result of a four-element vector over two or more passes of the data path 1300 during a single instruction cycle.
[0151] In another embodiment, Figure 13The logic shown can be replicated one or more times to produce multiple dot product results in parallel using shared elements from operand collector 920. For example, data path 1300 can include four four-element dot product logic units that match the logic shown in Figure 13 All four dot product logic units share the same vector during a particular pass, but are loaded with different vectors to produce four dot product results in parallel for different elements of collector matrix C 730. Additionally, during a given instruction cycle, multiple passes of data path 1300 can be used to produce dot product values for different vectors stored in operand collector 920.
[0152] In yet another embodiment, eight four-element dot product logic units that match the logic shown in Figure 13 can be included in the data path such that eight dot product values corresponding to two vectors and four vectors can be generated in a single pass. It should be appreciated that any number of copies of the logic shown in Figure 13 can be implemented in a single data path 1300 to produce multiple separate and distinct dot product values in parallel. Each copy of the logic can then be paired with any two four-element vectors stored in operand collector 920 to generate a dot product value for those two vectors. Each copy of the logic can also be coupled to selection logic that enables different pairs of vectors stored in the operand collector during a single instruction cycle to be consumed by that copy of the logic during multiple passes of the data path.
[0153] Figure 14 shows a double-precision floating-point FMA data path configured to share at least one pipeline stage with Figure 10 in accordance with one embodiment. Figure 13The HMMA data path 1300. It should be appreciated that when analyzing the architectures of the data path 1000, data path 1100, data path 1200, and data path 1300, the fourth pipeline stage 1304, fifth pipeline stage 1305, and sixth pipeline stage 1306 look relatively familiar, as described above. Basically, these pipeline stages all include a completion adder, normalization logic, and rounding logic to convert the carry value and sum value generated for the dot product to fit the specific format of the elements of the collector matrix C 730. The only difference between the logic in each of the data paths 1000, 1100, 1200, and 1300 is the precision of the logic. However, the completion adder 1050 of the double-precision floating-point FMA data path 1000 will be larger than the completion adders of the HMMA data paths 1100, 1200, and 1300. Therefore, the data path 1300 can be simplified in any architecture where the core includes both the HMMA data path 1300 and the double-precision floating-point FMA data path 1000, and the double-precision floating-point FMA data path 1000 is coupled to the same operand collector 920 and result queue 950.
[0154] In one embodiment, the core 450 includes both the HMMA data path 1300 and the double-precision floating-point FMA data path 1000. However, the HMMA data path 1300 is modified to omit the fourth pipeline stage 1304, fifth pipeline stage 1305, and sixth pipeline stage 1306. Instead, the output of the third pipeline stage 1303 (i.e., the carry value and sum value representing the dot product value output by the CSA tree) is routed to a pair of switches included in the FMA data path 1000. This pair of switches enables the FMA data path 1000 to sum either the dot product value from the HMMA data path 1300 or the FMA result from the FMA data path 1000 using the completion adder. Therefore, the HMMA data path 1300 shares the pipeline stages of the FMA data path 1000 that include the completion adder, normalization logic, and rounding logic. It should be appreciated that although not explicitly shown, the maximum exponent value associated with the dot product can also be sent to the switches in the FMA data path 1000 such that the normalization logic can convert between the exponent value generated by the adder 1020 and the maximum exponent value associated with the dot product generated by the HMMA data path 1300.
[0155] Sharing the logic between the two data paths, as well as sharing the operand collector 920 and result queue 950, can significantly reduce the size of the die footprint of the core 450 on the integrated circuit. Therefore, more cores 450 can be designed on a single integrated circuit die.
[0156] As described above, various data paths can be designed to implement MMA operations more efficiently than in current data path designs, such as scalar FMA data paths and even vector machines configured to compute partial products in parallel. A main aspect of this design is that more than one pair of vectors can be loaded from the register file and coupled to the input of the data path so that multiple dot product values can be generated in a single instruction cycle. As used herein, an instruction cycle refers to all operations related to loading an operand collector having multiple operands from the register file, then performing an MMA operation on the data path to generate multiple dot product values corresponding to different elements of the result matrix, and then writing the multiple dot product values to the register file. Each instruction cycle can include multiple passes of the data path to generate results for combinations of different vector pairs for the multiple passes. In addition, each pass can be pipelined such that the second pass starts before the first pass is complete. Instructions for MMA operations can be implemented over multiple instruction cycles, during each of which different portions of the input matrix operands are loaded into the operand collector of the data path. Thus, instructions for MMA operations can include matrix operands of any size processed over multiple instruction cycles and / or multiple cores, which during each instruction cycle apply different vectors from the matrix operand to each data path until all vectors from the matrix operand have been processed.
[0157] Known applications of MMA operations include image processing (e.g., performing an affine transformation on an image), machine learning (e.g., using matrix operations when performing linear algebra, optimization functions, or computing statistics), and others. Matrix algebra is a fundamental field that can be widely applied to various applications. Therefore, improving the processing efficiency of MMA operations by designing a processor capable of performing these operations faster is highly beneficial for the speed and efficiency of computational processing.
[0158] More specifically, MMA operations performed using the disclosed data paths exhibit better digital behavior and / or provide higher efficiency of the processor implementing the data path. For example, parallel accumulation of partial products using a single adder eliminates multiple rounding compared to using a serial adder that performs rounding each time as part of the accumulation operation. For vectors of any length, the worst-case error bound can be pushed to one (or half) unit of machine precision, while a serial multiply-add-add-add (mul-add-add-add) data path implemented in a conventional dot product data path shows a worst-case error bound proportional to the vector length.
[0159] In addition, the disclosed data path exhibits lower latency than prior art data paths. Fewer pipeline stages require fewer flip-flops to be implemented, which improves the power consumption of the data path. Since the data path reuses operand vectors, a smaller register file is needed to implement the same MMA operation, which can be achieved by serial operations of a traditional data path. In addition, the internal shifter and adder lengths can be simplified to match the desired error bounds, thereby further reducing the number of flip-flops in the data path. Furthermore, energy can be saved by simply updating some of the matrix operands at the operand collector at the input of the data path, chaining dot product operations to generate larger dot product results in the result queue without having to force the intermediate results to be written back to the register file and then reload the intermediate results from the register file back to the operand collector at the input of the data path.
[0160] Figure 15 An exemplary system 1500 is shown in which various architectures and / or functions of the various previous embodiments can be implemented. As shown, a system 1500 is provided that includes at least one central processor 1501 connected to a communication bus 1502. The communication bus 1502 can be implemented using any suitable protocol, such as PCI (Peripheral Component Interconnect), PCI-Express, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol. The system 1500 also includes a main memory 1504. Control logic (software) and data are stored in the main memory 1504, which can take the form of random access memory (RAM).
[0161] The system 1400 also includes an input device 1512, a graphics processor 1506, and a display 1508, i.e., a conventional CRT (Cathode Ray Tube), LCD (Liquid Crystal Display), LED (Light Emitting Diode), plasma display, etc. User input can be received from the input device 1512, such as a keyboard, mouse, touchpad, microphone, etc. In one embodiment, the graphics processor 1506 can include multiple shader modules, rasterization modules, etc. Each of the foregoing modules can even be located on a single semiconductor platform to form a graphics processing unit (GPU).
[0162] In this specification, a single semiconductor platform may refer to a unique single semiconductor-based integrated circuit or chip. It should be noted that the term single semiconductor platform may also refer to a multi-chip module with increased connectivity that emulates on-chip computing and represents a significant improvement over conventional central processing unit (CPU) and bus implementations. Of course, depending on the user's needs, various modules may also be positioned individually or in various combinations of semiconductor platforms.
[0163] System 1500 may also include auxiliary memory 1510. Auxiliary memory 1510 includes, for example, a hard disk drive and / or a removable storage drive, representing a floppy disk drive, a tape drive, an optical disk drive, a digital versatile disk (DVD) drive, a recording device, a universal serial bus (USB) flash drive. The removable storage drive reads from and / or writes to a removable storage unit in a well-known manner.
[0164] A computer program or computer control logic algorithm may be stored in main memory 1504 and / or auxiliary memory 1510. These computer programs, when executed, enable system 1500 to perform various functions. Memory 1504, memory 1510, and / or any other memory are possible examples of computer-readable media.
[0165] In one embodiment, the architecture and / or functionality of the various previous figures may be implemented in a central processing unit 1501, a graphics processing unit 1506, an integrated circuit (not shown) having at least a portion of the capabilities of both the central processing unit 1501 and the graphics processing unit 1506, a chipset (i.e., a group of integrated circuits designed to work and be sold as a unit for performing related functions), and / or any other integrated circuit for this purpose.
[0166] Still, the architecture and / or functionality of the various previous figures may be implemented in a general-purpose computer system, a circuit board system, a game console system dedicated to entertainment purposes, a dedicated system, and / or any other desired system. For example, system 1500 may take the form of a desktop computer, a laptop computer, a server, a workstation, a game console, an embedded system, and / or any other type of logic. Additionally, system 1500 may take the form of various other devices, including but not limited to personal digital assistant (PDA) devices, mobile phone devices, televisions, etc.
[0167] Furthermore, although not shown, system 1500 for communication purposes may be coupled to a network (e.g., a telecommunications network, a local area network (LAN), a wireless network, a wide area network (WAN) such as the Internet, a peer-to-peer network, a cable television network, etc.).
[0168] Although various embodiments have been described above, it should be understood that they are presented by way of example only and not by way of limitation. Accordingly, the breadth and scope of the preferred embodiments should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the claims and their equivalents.
Claims
1. A multi-threaded processor, comprising: A decoder for decoding matrix multiply-accumulate instructions for signed matrix data; A buffer for storing the signed matrix data specified by the operands of the matrix multiply-accumulate instructions; A scheduler for scheduling the matrix multiply-accumulate instructions; A fused multiply-accumulate unit for performing a dot product of corresponding elements of the signed matrix data; An arithmetic logic unit for adding multiple partial product results of the dot product to be accumulated into a register; And A memory for storing the result of the matrix multiply-accumulate instructions.
2. The multi-threaded processor according to claim 1, wherein the signed matrix data comprises 32-bit two's complement integer data.
3. The multi-threaded processor according to claim 1, wherein the signed matrix data comprises 16-bit two's complement integer data.
4. The multi-threaded processor according to claim 1, further comprising an adder tree, wherein the adder tree comprises at least 3:2 carry-save adders.
5. The multi-threaded processor according to claim 1, further comprising a dispatch unit for sending the matrix multiply-accumulate instruction to the fused multiply-accumulate unit.
6. The multi-threaded processor according to claim 1, further comprising a register file for providing the registers.
7. The multi-threaded processor according to claim 1, further comprising a special function unit.
8. The multi-threaded processor according to claim 1, further comprising an interconnect for connecting the arithmetic logic unit to the registers.
9. A system comprising the multi-threaded processor according to claim 1, wherein the system further comprises: A system bus for connecting the multi-threaded processor to one or more peripheral devices; And One or more dynamic random access memory devices.
10. A single instruction multiple data multi-threaded processor, comprising: A plurality of cores for executing matrix multiply-accumulate instructions, wherein each core of the plurality of cores includes: A front end for fetching the matrix multiply-accumulate instructions; An instruction cache for storing the matrix multiply-accumulate instructions; An L1 cache for storing data; An L2 cache for storing data; A plurality of ports for reading from and writing to memory; One or more load or store units for reading from and writing to the memory; An interconnect for coupling the memory and the plurality of cores; A decoder for decoding the matrix multiply-accumulate instructions; A buffer for storing the signed matrix data specified by the operands of the matrix multiply-accumulate instructions; A scheduler for scheduling the matrix multiply-accumulate instructions; A fused multiply-accumulate unit for performing a dot product of corresponding elements of the signed matrix data; An arithmetic logic unit for adding multiple partial product results of the dot product to be accumulated into a register; and Wherein the memory is used to store the result of the matrix multiply-accumulate instructions.
11. The single instruction multiple data multi-threaded processor according to claim 10, wherein the L1 cache includes at least 24 kilobytes of storage.
12. The single instruction multiple data multi-threaded processor according to claim 10, wherein the memory includes at least 64 kilobytes of storage.
13. The single instruction multiple data multi-threaded processor according to claim 10, wherein the interconnect connects the one or more load or store units to the register.
14. The single instruction multiple data multi-threaded processor according to claim 10, wherein the scheduler dispatches the matrix multiply and accumulate instructions to one or more of the plurality of cores.
15. The single instruction multiple data multi-threaded processor according to claim 10, wherein the signed matrix data includes 32-bit two's complement integer data.
16. The single instruction multiple data multi-threaded processor according to claim 10, wherein the signed matrix data includes 16-bit two's complement integer data.
17. A computer-implemented method, comprising: Decoding matrix multiply-accumulate instructions for signed matrix data by the decoder; Storing the signed matrix data specified by the operands of the matrix multiply-accumulate instructions by the buffer; Scheduling the matrix multiply-accumulate instructions by the scheduler; Performing a dot product of corresponding elements of the signed matrix data by the fused multiply-accumulate unit; Adding multiple partial product results of the dot product to be accumulated into a register by the arithmetic logic unit; And Storing the result of the matrix multiply-accumulate instructions by the memory.
18. The computer-implemented method according to claim 17, wherein the arithmetic logic unit includes at least one adder.
19. The computer-implemented method according to claim 17, wherein the signed matrix data includes 32-bit two's complement integer data.
20. The computer-implemented method according to claim 17, wherein the signed matrix data includes 16-bit two's complement integer data.
21. The computer-implemented method according to claim 17, wherein the scheduler includes a dispatch unit for dispatching the matrix multiply and accumulate instructions.
22. The computer-implemented method according to claim 17, further comprising using an interconnect to add the plurality of partial accumulations to the register.
23. The computer-implemented method according to claim 17, wherein the register file provides the registers.
Citation Information
Patent Citations
X87 fused multiplication-addition instruction and its use
CN101145099A
Calculation engine and electronic equipment
CN106126481A