Stream processor with low-power parallel matrix multiplication pipeline

By designing a low-power parallel matrix multiplication pipeline in a stream processor, using the coordinated work of multiple dot product units and vector register stacks, the problems of high power consumption and long waiting time in the prior art are solved, and efficient matrix multiplication calculation is realized.

CN109871236BActive Publication Date: 2025-05-06ADVANCED MICRO DEVICES INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201711249532.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2017-12-01
Publication Date
2025-05-06
Estimated Expiration
2037-12-01

AI Technical Summary

Technical Problem

Existing stream processors consume high power and have long wait times when performing matrix multiplication operations, which affects computing performance.

Method used

A low-power parallel matrix multiplication pipeline is designed, using execution pipelines of multiple dot product units, and the efficient execution of matrix multiplication operations is achieved through the coordinated work of vector register stack and execution pipeline.

Benefits of technology

This design significantly reduces the power consumption of matrix multiplication operations and reduces latency, improving the computing performance of the stream processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN109871236B_ABST
    Figure CN109871236B_ABST
Patent Text Reader

Abstract

The present invention relates to a stream processor with a low-power parallel matrix multiplication pipeline. A system, an apparatus and a method for implementing a low-power parallel matrix multiplication pipeline are disclosed. In one embodiment, the system includes at least a first and a second vector register file coupled to the matrix multiplication pipeline. The matrix multiplication pipeline includes a plurality of dot product units. The dot product unit is configured to calculate the dot product or outer product of the first and second groups of operands retrieved from the first vector register file. The result of the dot product or outer product operation is written back to the second vector register file. The second vector register file uses the result of the previous dot product or outer product operation as the input of the subsequent dot product or outer product operation. The dot product unit receives the results from the previous stages of the matrix multiplication operation and accumulates the results of these previous dot products or outer products with the current dot product or outer product result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and more particularly to a stream processor with a low-power parallel matrix multiplication pipeline. Background Art

[0002] Many different types of computing systems include vector processors or single instruction multiple data (SIMD) processors. Tasks can be executed in parallel on these types of parallel processors to increase the throughput of the computing system. Note that parallel processors may also be referred to as "stream processors" here. Various types of machine learning algorithms are implemented on stream processors. Some of these machine learning algorithms implement matrix multiplication operations. These matrix multiplication operations typically require many cycles to produce results while consuming a lot of power. Therefore, it is desirable to use techniques for improving performance, reducing power consumption, and / or reducing the latency of matrix multiplication operations on stream processors. Summary of the invention

[0003] Some aspects of the present invention are embodied as follows:

[0004] 1. A system comprising: a first vector register file; and a first execution pipeline coupled to the first vector register file, wherein the first execution pipeline comprises a plurality of dot product units, and wherein each of the plurality of dot product units is configured to: calculate a plurality of products of elements in a first set of operands and corresponding elements in a second set of operands; and calculate a sum of an accumulated input and the plurality of products, wherein the sum is an output of the dot product unit.

[0005] 2. The system of clause 1, wherein the system is configured to read the first set of operands and the second set of operands from the first vector register file and provide the first set of operands and the second set of operands to the first execution pipeline.

[0006] 3. The system of clause 1, wherein the system further comprises a second vector register file, wherein the system is configured to read a plurality of accumulated inputs from the second vector register file and provide the plurality of accumulated inputs to the first execution pipeline.

[0007] 4. The system of clause 3, wherein each dot product unit is further configured to write the output to the second vector register file, wherein the output of a previous dot product operation is an accumulated input added to the sum of the current dot product operation.

[0008] 5. The system of clause 1, wherein the system further comprises a second execution pipeline, wherein the second execution pipeline is configured to perform operations in parallel with the dot product operation being performed by the first execution pipeline.

[0009] 6. A system as described in clause 5, wherein the system is further configured to: read the first group of operands and the second group of operands from the first vector register file in a first cycle and store the first group of operands and the second group of operands in a storage element; and read a third group of operands from the first vector register file in a second cycle and provide the third group of operands to the second execution pipeline.

[0010] 7. A system as described in clause 1, wherein the first set of operands are rows of a first matrix, and wherein the second set of operands are columns of a second matrix, and wherein the plurality of dot product units are configured to multiply the first matrix by the second matrix.

[0011] 8. A method comprising: computing a plurality of products of elements of a first set of operands and corresponding elements of a second set of operands; and computing a sum of an accumulated input and the plurality of products, wherein the sum is an output of a dot product unit.

[0012] 9. The method of clause 8, further comprising reading the first set of operands and the second set of operands from a first vector register file and providing the first set of operands and the second set of operands to a first execution pipeline.

[0013] 10. The method of clause 8, further comprising reading a plurality of accumulated inputs from a second vector register file and providing the plurality of accumulated inputs to the first execution pipeline.

[0014] 11. The method of clause 10, further comprising writing the output to the second vector register file, wherein the output of a previous dot product operation is an accumulated input added to a sum of a current dot product operation.

[0015] 12. The method of clause 8, further comprising performing an operation on the second execution pipeline in parallel with the dot product operation being performed by the first execution pipeline.

[0016] 13. The method as described in clause 12 also includes: reading the first group of operands and the second group of operands from a first vector register file in a first cycle and storing the first group of operands and the second group of operands in a storage element; and reading a third group of operands from the first vector register file in a second cycle and providing the third group of operands to the second execution pipeline.

[0017] 14. The method of clause 8, wherein the first set of operands are rows of a first matrix, and wherein the second set of operands are columns of a second matrix, and wherein the method further comprises multiplying the first matrix by the second matrix using a plurality of dot product units.

[0018] 15. An apparatus comprising: a plurality of vector register files; and a plurality of execution pipelines coupled to the plurality of vector register files; wherein the apparatus is configured to: compute a plurality of products of elements of a first set of operands and corresponding elements of a second set of operands; and compute a sum of an accumulated input and the plurality of products, wherein the sum is an output of a dot product unit.

[0019] 16. The apparatus of clause 15, wherein the apparatus is further configured to read the first set of operands and the second set of operands from a first vector register file and provide the first set of operands and the second set of operands to a first execution pipeline.

[0020] 17. An apparatus as recited in clause 16, wherein the apparatus is configured to read a plurality of accumulated inputs from a second vector register file and provide the plurality of accumulated inputs to the first execution pipeline.

[0021] 18. The apparatus of clause 17, wherein the apparatus is configured to write the output to the second vector register file, wherein an output of a previous dot product operation is an accumulated input added to a sum of a current dot product operation.

[0022] 19. The apparatus of clause 15, wherein the apparatus is further configured to perform an operation on the second execution pipeline in parallel with a dot product operation being performed by the first execution pipeline.

[0023] 20. An apparatus as described in clause 19, wherein the apparatus is further configured to: read the first group of operands and the second group of operands from the first vector register file in a first cycle and store the first group of operands and the second group of operands in a storage element; and read a third group of operands from the first vector register file in a second cycle and provide the third group of operands to the second execution pipeline. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Advantages of the methods and mechanisms described herein may be better understood by referring to the following description in conjunction with the accompanying drawings, in which:

[0025] Figure 1 is a block diagram of one implementation of a computing system.

[0026] Figure 2 is a block diagram of one implementation of a matrix multiplication operation.

[0027] Figure 3 is a block diagram of one implementation of a stream processor.

[0028] Figure 4 is a timing diagram of one implementation of overlapping execution on an execution pipeline.

[0029] Figure 5 is a timing diagram of another embodiment of overlapping execution on an execution pipeline.

[0030] Figure 6 is a block diagram of another implementation of a matrix multiplication operation.

[0031] Figure 7 is a block diagram of another implementation of a stream processor.

[0032] Figure 8 is a timing diagram for one embodiment of performing a matrix multiplication operation.

[0033] Fig. 9 is a timing diagram of another implementation for performing a matrix multiplication operation.

[0034] Fig.10 is a block diagram of another implementation of a matrix multiplication operation.

[0035] Fig.11 is a block diagram of another implementation of a stream processor.

[0036] Fig.12 is a timing diagram for one embodiment of performing a matrix multiplication operation.

[0037] Fig.13 is a timing diagram of another implementation for performing a matrix multiplication operation.

[0038] Fig.14 is a generalized flow chart illustrating one embodiment of a method for performing a matrix multiplication operation. DETAILED DESCRIPTION

[0039] In the following description, many specific details are set forth to provide a thorough understanding of the methods and mechanisms presented herein. However, it will be appreciated by those of ordinary skill in the art that various embodiments may be practiced without these specific details. In some cases, known structures, components, signals, computer program instructions, and techniques are not shown in detail to avoid blurring the methods described herein. It should be understood that, for simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the size of some elements may be enlarged relative to other elements.

[0040] Disclosed herein are systems, devices, and methods for implementing a low-power parallel matrix multiplication pipeline. In one embodiment, a stream processor includes a plurality of vector register files and a plurality of execution pipelines coupled to the vector register files. The first execution pipeline includes a plurality of dot product units. In one embodiment, each of these dot product units is configured to perform a dot product operation on the first and second groups of operands by calculating the sum of a plurality of products of elements of the first group of operands and corresponding elements of the second group of operands. Each dot product unit is also configured to generate an output equal to an accumulated value added to the result of the dot product operation. In one embodiment, the accumulated value is the result of a previous dot product operation. In another embodiment, each dot product unit is configured to perform a matrix multiplication operation by calculating the outer product of the first and second groups of operands.

[0041] In one embodiment, the stream processor is configured to read the first and second sets of operands from the first vector register file and provide the first and second sets of operands to the first execution pipeline. In this embodiment, the stream processor is configured to read a plurality of accumulated values ​​from the second vector register file and provide the plurality of accumulated values ​​to the first execution pipeline. Furthermore, the first execution pipeline is configured to write the output generated by the dot product unit to the second vector register file.

[0042] Reference now Figure 1 , a block diagram of one embodiment of a computing system 100 is shown. In one embodiment, the computing system 100 includes at least a processor 110, an input / output (I / O) interface 120, a bus 125, and a memory device 130. In other embodiments, the computing system 100 may include other components and / or the computing system 100 may be arranged differently. The processor 110 represents any number and type of processing units (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC)).

[0043] In one embodiment, the processor 110 includes a vector processor having a plurality of stream processors 115. Each stream processor 115 may also be referred to as a processor or a processing channel. In one embodiment, each stream processor 115 includes at least two types of execution pipelines (e.g., matrix multiplication pipelines, fused multiply-add (FMA) pipelines) that share one or more vector register files. In one embodiment, each vector register file includes a multi-bank high-density random access memory (RAM). In various embodiments, the execution of instructions may be overlapped on multiple execution pipelines to increase the throughput of the stream processor.

[0044] The memory device 130 represents any number and type of memory devices. For example, the memory types in the memory device 130 may include dynamic random access memory (DRAM), static random access memory (SRAM), NAND flash memory, NOR flash memory, ferroelectric random access memory (FeRAM), etc. The memory device 130 can be accessed by the processor 110. The I / O interface 120 represents any number and type of I / O interface (e.g., peripheral component interconnect (PCI) bus, PCI-extension (PCI-X), PCIE (PCI Express) bus, Gigabit Ethernet (GBE) bus, universal serial bus (USB)). Various types of peripheral devices can be coupled to the I / O interface 120. These peripheral devices include (but are not limited to) a display, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, and the like.

[0045] In various embodiments, computing system 100 may be any of a computer, a laptop, a mobile device, a server, a game console, or various other types of computing systems or devices. Note that the number of components of computing system 100 may vary depending on the embodiment. Each component / subcomponent may be larger than Figure 2 The number shown in FIG. 1 may be greater or less than that shown in FIG. 1 . Note also that computing system 100 may include Figure 1 Other components not shown.

[0046] Now go to Figure 2 , a block diagram 200 of one embodiment of a matrix multiplication operation is shown. In one embodiment, matrix 202 is multiplied by matrix 204 to generate matrix 206. Matrix 202 may also be referred to as matrix A, matrix 204 may also be referred to as matrix B, and matrix 206 may also be referred to as matrix C. In one embodiment, matrix 202 is a 32×4 matrix and matrix 204 is a 4×32 matrix. Matrix 202 and matrix 204 may be stored in any bank of a vector general purpose register (VGPR) stack. In some embodiments, matrix 202 is part of a first matrix and matrix 204 is part of a second matrix. The first matrix and the second matrix may be split into smaller matrices, and the matrix multiplication operation is performed on the smaller matrices.

[0047] In one embodiment, the data for each entry in matrix 202 and matrix 204 is a 16-bit floating point value. In other embodiments, the data may be represented in other formats and / or other numbers of bits. In one embodiment, matrix 202 includes the values ​​of an input data set, and matrix 204 includes weighted values ​​to be applied to the input data set. In this embodiment, the elements of the input data set are multiplied by the weighted values ​​and then accumulated into the sum of the neurons representing the neural network. In one embodiment, the neuron may be compared to a threshold value to determine whether the neuron is activated by the input value. In other embodiments, other types of decisions may be made based on the neuron value, and / or the neuron value may be fed to another layer of the neural network.

[0048] In one embodiment, an outer product matrix multiplication operation is performed on matrix 202 and matrix 204 to generate matrix 206. The outer product matrix multiplication operation is performed to minimize the bandwidth of the internal and external memory used when extracting the input matrices 202 and 204. The outer product matrix multiplication operation also reduces the data movement through the processor. For example, in one embodiment, the elements of matrices 202 and 204 are extracted once and then reused in multiple cycles. Moreover, in one embodiment, when matrix 204 is provided to the matrix multiplication pipeline, data path switching is reduced by keeping matrix 204 unchanged.

[0049] like Figure 2 As shown in FIG. 200 , during a first cycle (cycle 0), a matrix multiplication pipeline is used to multiply the first row of matrix 202 by the column of matrix 204. It should be noted that a matrix multiplication pipeline may also be referred to as a matrix multiplication unit. In a second cycle (cycle 1), the second row of matrix 202 is multiplied by the column of matrix 204. This pattern may continue during the remaining cycles of the 32 cycles to complete the matrix multiplication operation between matrix 202 and matrix 204 to produce matrix 206.

[0050] In one embodiment, a plurality of four-operand dot product (sometimes referred to as inner product) operations are performed between a first row of matrix 202 and a column of matrix 204 in a first clock cycle. Then, a plurality of four-operand dot product operations are performed between a second row of matrix 202 and a column of matrix 204 in a second clock cycle. This pattern may continue for the remaining rows of matrix 202 for the other cycles in the 32 cycle sequence. In another embodiment, a matrix multiplication operation is performed in a first clock cycle by calculating the outer product of the first row of matrix 202 and the column of matrix 204. In a second clock cycle, the outer product of the second row of matrix 202 and the column of matrix 204 is calculated. This pattern is continued for the other rows of matrix 202. Note that in other embodiments, the size of the matrix and / or the size of the matrix multiplication pipeline may vary.

[0051] Reference now Figure 3 , a block diagram of an embodiment of a stream processor 300 is shown. In one embodiment, components of the stream processor 300 are included in each stream processor 115 ( Figure 1 ). It should be noted that the architecture of the stream processor 300 is intended to represent a specific implementation of the stream processor. It should be understood that in other embodiments, the architecture of the stream processor 300 may vary. For example, the data width of some paths is indicated in the entire architecture (e.g., 128 bits (b), 32b), but in other embodiments, these paths may have other widths. Moreover, the number of channels per path may also be different from the number of channels shown in the stream processor 300. In addition, although 32 DOT4 units 330A-H are shown in the matrix multiplication pipeline of the stream processor 300, other pipelines may also have other numbers and / or other sizes of dot product units (e.g., DOT8 units).

[0052] In one embodiment, stream processor 300 includes two separate vector register files 304 and 308. Vector register files 304 and 308 may also be referred to as vector general purpose register (VGPR) files. In addition, VGPR file 304 may be referred to as accumulation VGPR file 304, and VGPR file 308 may be referred to as architecture VGPR file 308. Accumulation VGPR file 304 and source multiplexer 310 are coupled together to build a single VGPR file, which provides multiple read ports X, Y, Z and W. Therefore, matrix C and matrix D can be stored in any library of accumulation VGPR file 304. Architecture VGPR file 308 and source multiplexer 312 are coupled together to build a single VGPR file, which provides multiple read ports A, B, C and D. Therefore, matrix A and matrix B can be stored in any library of architecture VGPR file 308.

[0053] In one embodiment, the outputs of DOT4 cells 330A-H are coupled back to the inputs of accumulator VGPR 304 via multiplexer (or mux) 302. The operands of source X and source Y read from banks 0 and 1 of accumulator VGPR 304 are coupled to the inputs of DOT4 cells 330A-H via multiplexer 310. Also, the operands of source Z and source W are coupled to accumulator VGPR export unit 314 to be written to memory (not shown) or another location. In one embodiment, each DOT4 cell 330A-H is configured to generate a dot product of two input vectors. For example, for input vectors X and Y having elements i from 0 to 3, the dot product generated by each DOT4 cell 330A-H is equal to x 0 y 0 +x 1 y 1 +x2 y 2 +x 3 y 3 Each DOT4 unit 330A-H may also add intermediate results to the dot product, thereby computing a longer dot product by performing multiple quaternary dot products and accumulating the intermediate results. For example, the dot product of the (i+1) iteration may be computed by each DOT4 unit 330A-H as: dot-product(i+1)=x 0 y 0 +x 1 y 1 +x 2 y 2 +x 3 y 3 +dot-product(i). Each DOT4 unit 330A-H includes a plurality of multiplier accumulators (MACs) to perform dot product operations. In another embodiment, each DOT4 unit 330A-H is configured to generate the outer product of two input vectors. For example, for each input vector having four elements, the outer product generated by each DOT4 unit 330A-H will be a 4×4 matrix.

[0054] As described above, a first set of operands is coupled to the DOT4 units 330A-H from the accumulation VGPR 304. Furthermore, a second set of operands is coupled to the DOT4 units 330A-H from the architectural VGPR 308. The second set of operands includes elements of the A and B matrices read from banks 0 to 3 of the VGPR 308. The intermediate results of the matrix multiplication operation of the A and B matrices are written to the accumulation VGPR 304, and the intermediate results are routed back to the DOT4 units 330A-H from banks 0-1 of the accumulation VGPR 304. In addition, operands from bank 2 of the architectural VGPR 308 are coupled to the FMA pipeline 324 and the vector input / output (I / O) derivation unit 318. Operands from bank 3 of the architectural VGPR 308 are coupled to the vector input / output (I / O) derivation unit 318. The four banks of the architectural VGPR 308 are used to implement a pseudo multi-port register file. Source multiplexer 312 is designed to provide this multi-port capability to architectural VGPR 308. The output from FMA pipeline 324 is coupled back to architectural VGPR 308 through multiplexer 306. Note that in other embodiments, accumulation VGPR 304 and architectural VGPR 308 may have other numbers of banks besides four.

[0055] In one embodiment, source A and B operands are coupled from the architectural VGPR 308 to the DOT4 units 330A-H via a data path having multiple components. In one embodiment, these data paths include a source multiplexer 312, an architecture register rotation crossbar 316, double buffers 320 and 322, and a crossbar 326. The architecture register rotation crossbar 316 is used to rotate the A and B operands into the appropriate lanes to couple to the DOT4 units 330A-H to perform dot product operations on the appropriate matrix elements. The double buffer 320 for the A operand and the double buffer 322 for the B operand are used to store the operands so that the operands can be used in multiple cycles without having to be re-fetched from the architectural VGPR 308. The output of the double buffer 320 is coupled to a 4×4 matrix replication crossbar 326 to rotate the operands between lanes depending on the stage of the matrix multiplication operation being performed. Note that in other implementations, other suitable types of buffers may be used instead of double buffers 320 and 322 .

[0056] In one embodiment, operands are coupled to the DOT4 units 330A-H from the accumulation VGPR 304 and the architectural VGPR 308 to reduce external memory bandwidth utilization of the stream processor 300 when performing matrix multiplication operations. The elements of the A and B matrices are read a single time from the architectural VGPR 308 and then fed to the DOT4 units 330A-H from the double buffers 320 and 322 over multiple cycles. In one embodiment, the elements of the B matrix coupled to the DOT4 units 330A-H are not toggled during these multiplication cycles. This helps to reduce the amount of power consumed in the matrix multiplication operation.

[0057] The A and B operands from the architectural VGPR 308 are also coupled to the fused multiply-add (FMA) pipeline 324. When the A and B operands are read from the architectural VGPR 308 and coupled to the DOT4 units 330A-H in a first clock cycle, these A and B operands can be reused in subsequent clock cycles. This allows operands to be read from the architectural VGPR 308 in subsequent clock cycles and provided to the FMA pipeline 324. This enables overlapping concurrent execution to occur on the pipelines 330 and 324.

[0058] Now go to Figure 4 , shows a timing diagram 400A of one embodiment of overlapping execution on an execution pipeline. For the purpose of discussion, it can be assumed that the timing diagram 400A is applicable to ( Figure 3400A) stream processor 300. The operations shown in timing diagram 400 are merely indicative of a particular implementation. In other implementations, other sequences of operations may be performed on stream processor 300. The cycles shown at the top of timing diagram 400A indicate clock cycles of stream processor 300. In one implementation, each cycle shown in timing diagram 400A represents four actual clock cycles for the dot product unit of the matrix multiplication pipeline to produce the result of a given stage of the matrix multiplication operation. In other implementations, each cycle shown in timing diagram 400A may represent other numbers of actual clock cycles.

[0059] In cycle 0, operands of source A and source B are read from the architectural VGPR stack, and operands of source X and source Y are read from the accumulation VGPR stack. These operands are provided to the matrix multiplication pipeline to be used in cycle 1. During cycle 1, source operands can be read from the architectural VGPR stack and provided to the FMA pipeline so that execution can overlap on both the matrix multiplication pipeline and the FMA pipeline. This allows the stream processor to perform different operations in concurrent cycles. Moreover, during cycle 1, operands of source X and Y are read from the accumulation VGPR stack and provided to the matrix multiplication pipeline to be used in cycle 2. This pattern can continue for subsequent cycles, with operands of source X and Y being read from the accumulation VGPR stack and provided to the matrix multiplication pipeline.

[0060] Also, during cycle 1, the accumulation source Z operands may be read from the accumulation VGPR file. These accumulation source Z operands may then be written to memory in cycle 2. This pattern of reading accumulation source Z operands from the accumulation VGPR file and then writing these values ​​to memory may occur in subsequent cycles. Also, the operands for sources A and B may be stored in a double buffer (or other temporary storage) and rotated in subsequent cycles to shift the operands to the appropriate lanes of the matrix multiplication pipeline.

[0061] Reference now Figure 5 , shows another embodiment of a timing diagram 400B for overlapping execution on an execution pipeline. The timing diagram 400B is intended to represent ( Figure 4 In subsequent cycles 6-10, the same mode of operation shown in timing diagram 400A may be performed for the ( Figure 3 The timing diagram 400B of the operations performed on the stream processor 300 continues.

[0062] In one embodiment, in cycle 8, a matrix multiplication operation is completed for the first set of matrix elements. In cycle 8, a new set of matrix elements is retrieved from the architectural VGPR stack and read into the double buffer. During cycle 9, since the FMA pipeline will not be able to access the architectural VGPR stack during cycle 8, there is a bubble in the FMA pipeline. However, starting from cycle 9, the FMA pipeline can access the architectural VGPR stack again and begin reading operands for new FMA operations, which can be performed in parallel with the matrix multiplication operations performed in cycle 10 and subsequent cycles. When Figure 400B stops in cycle 10, subsequent cycles can follow the same operating mode shown in Figures 400A-B.

[0063] Now go to Figure 6 , a block diagram 600 of another embodiment of a matrix multiplication operation is shown. In one embodiment, an A matrix 602 is multiplied by a B matrix 604 to generate a C matrix 606. In one embodiment, the A matrix 602 is divided into 4×8 portions, and the B matrix 604 is divided into 8×4 portions to perform the matrix multiplication operation. As shown in the diagram 600, in cycle 0, the first row of the A matrix 602 is multiplied by each column of the B matrix 604 to generate the first row of the C matrix 606. In cycle 1, the second row of the A matrix 602 is multiplied by each column of the B matrix 604 to generate the second row of the C matrix 606. For cycles 2-15, this pattern can continue for the remaining rows of the A matrix 602.

[0064] Reference now Figure 7 , a block diagram of another embodiment of the stream processor 700 is shown. In one embodiment, the components of the stream processor 700 include: Figure 1 ). In one embodiment, stream processor 700 includes two separate vector register files. The first vector register file is accumulation VGPR file 704. The output of DOT8 unit 730A-H is coupled back to the input of accumulation VGPR file 704 via multiplexer 702. The second register file is architectural VGPR file 708. The output of FMA pipeline 724 is coupled to the input of architectural VGPR file 708 via multiplexer 706. Read ports X and Y of accumulation VGPR file 704 are coupled to input ports of DOT8 unit 730A-H through source multiplexer 710. Read ports Z and W of accumulation VGPR file 704 are coupled to derivation unit 714.

[0065] The DOT8 units 730A-H represent a matrix multiplication pipeline. In other embodiments, other numbers of DOT8 units may be combined to form matrix multiplication pipelines of other dimensions. For example, in another embodiment, 16 DOT8 units may be combined together to form a matrix multiplication pipeline. In another embodiment, 32 DOT8 units may be combined together to form a matrix multiplication pipeline. Other embodiments may include other numbers of DOT8 units. Moreover, in other embodiments, dot product units of other sizes (e.g., DOT4 units, DOT16 units) may be combined together and used to implement a matrix multiplication pipeline.

[0066] In one embodiment, each DOT8 cell 730A-H is configured to receive a signal from a second matrix (e.g., Figure 6 The corresponding eight elements of the B matrix 604) are implemented from the first matrix (eg, Figure 6 602) to generate a single output. These outputs are written back to the accumulation VGPR stack 704 and are also coupled back to the DOT8 units 730A-H to be added back to the next dot product operation performed for each subsequent group of eight elements from the first matrix and the corresponding eight elements from the second matrix. In another embodiment, each DOT8 unit 730A-H is configured to implement the dot product operation from the first matrix (e.g., Figure 6 The eight elements of the A matrix 602) are combined with the eight elements from the second matrix (e.g., Figure 6 The outer product operation of the corresponding eight elements of the B matrix 604) is performed to generate an 8×8 matrix.

[0067] The operands of ports A, B, and C of the architectural VGPR file 708 are coupled to the source multiplexer 712 and then pass through the crossbar 716. The operands of ports C and D of the architectural VGPR file 708 are both coupled to the vector I / O output unit 718. After the crossbar 716, the operands of ports A, B, and C of the architectural VGPR file 708 are coupled to double buffers 720, 722, and 723, respectively. The double buffers 720, 722, and 723 are configured to provide operands to the DOT8 units 730A-H for multiple cycles without having to read operands from the architectural VGPR file 708 in subsequent cycles. Therefore, operands can be read from ports A, B, and C of the architectural VGPR file 708 in one cycle and then used in multiple subsequent cycles. During these subsequent cycles, operands can be read from the architectural VGPR file 708 and provided to the FMA pipeline 724. This allows overlapping execution of different operations to occur after the first cycle on DOT8 units 730A-H and FMA pipeline 724. The output of FMA pipeline 724 is coupled back to architectural VGPR file 708 via multiplexer 706.

[0068] In one embodiment, operands from port C of the architecture VGPR stack 708 are coupled to DOT8 units 730E-H to be used in matrix multiplication operations. In this embodiment, operands from port B of the architecture VGPR stack 708 are coupled to DOT8 units 730A-D for matrix multiplication operations. Also, in this embodiment, operands from port A of the architecture VGPR stack 708 are coupled to DOT8 units 730A-H for matrix multiplication operations. In addition, operands from port A of the architecture VGPR stack 708 pass through a crossbar 726 to allow operands to be rotated to the correct channel for each stage of the matrix multiplication operation.

[0069] Now go to Figure 8 , shows one embodiment of a timing diagram 800A for performing a matrix multiplication operation. The timing diagram 800A is intended to represent Figure 7The timing of the operation of the stream processor 700 of the embodiment of the present invention. In one embodiment, in cycle 0, the operands of source A and source B from the architectural VGPR stack are read and coupled to the first matrix multiplication pipeline (i.e., DOT8 units 730A-D) of the stream processor. Moreover, in cycle 0, the operands of source A and source C from the architectural VGPR stack are read and coupled to the second matrix multiplication pipeline (i.e., DOT8 units 730E-H) of the stream processor. These operands for sources A, B, and C read from the architectural VGPR stack in cycle 0 are stored in a temporary memory (e.g., a double buffer) and reused in subsequent cycles. This helps to reduce the number of accesses to the architectural VGPR stack in subsequent cycles. In addition, this allows the FMA pipeline to fetch operands from the architectural VGPR stack in subsequent cycles and enables overlapping execution of the matrix multiplication pipeline and the FMA pipeline starting from cycle 2. Also, in cycle 0, operands for source X are read from the accumulation VGPR file and coupled to the first matrix multiplication pipeline, and operands for source Y are read from the accumulation VGPR file and coupled to the second matrix multiplication pipeline. In cycle 1, the matrix multiplication pipeline generates a dot product or outer product result of the first row of the output matrix C.

[0070] In cycle 1, source A, B, and C operands may be read from the architectural VGPR file and then coupled to the FMA pipeline in cycle 2. Also, in cycle 1, operands for sources X and Y may be read from the accumulation VGPR file and then provided to the first and second matrix multiplication pipelines, respectively, in cycle 2. Additionally, in cycle 1, the operand for source Z may be read from the accumulation VGPR file and then written to memory in cycle 2. As shown in timing diagram 800A, this mode of operation may continue for subsequent cycles 3-5. In subsequent cycles, the matrix multiplication pipeline generates subsequent rows in the output C matrix.

[0071] Reference now Fig. 9 , another embodiment of a timing diagram 800B for performing a matrix multiplication operation is shown. The timing diagram 800B is intended to be a sequential diagram of the timing diagram 800A ( Figure 8 ) is a continuation of the operations shown in ). In cycles 6, 7, and 8, the matrix multiplication pipeline generates additional rows in the output C matrix in the same pattern as shown in timing diagram 800A.

[0072] Now go to Fig.10, a block diagram of another embodiment of a matrix multiplication operation 1000 is shown. In one embodiment, an A matrix 1002 of size 16×8 is multiplied by a B matrix 1004 of size 8×16 to generate a C matrix 1006 of size 16×16. In one embodiment, the A matrix 1002 is multiplied by the B matrix 1004 using a matrix multiplication pipeline, the matrix multiplication pipeline including a dot product unit configured to perform a dot product or outer product operation on eight pairs of input operands. In one embodiment, the matrix multiplication operation of the A matrix 1002 multiplied by the B matrix 1004 requires 16 cycles.

[0073] Reference now Fig.11 , a block diagram of another embodiment of the stream processor 1100 is shown. In one embodiment, components of the stream processor 1100 are included in each stream processor 115 ( Figure 1 In one embodiment, the stream processor 1100 is configured to execute the graph 1000 ( Fig.10 ) as shown in the matrix multiplication operation. In one embodiment, the stream processor 1100 includes a single architectural VGPR file 1108. Figure 3 and Figure 7 Compared to the other stream processors 300 and 700 shown in FIG. 1 , stream processor 1100 does not include an accumulation VGPR file. Instead, the outputs of DOT8 units 1130A-H are coupled back to the architectural VGPR file 1108 via multiplexers 1106. Also, the outputs of FMA pipeline 1124 are coupled back to the architectural VGPR file 1108 via multiplexers 1106.

[0074] In one embodiment, the A matrix 1002 ( Fig.10 ) is stored in bank 0 of the architectural VGPR stack 1108, and the B matrix 1004 ( Fig.10 ) are stored in bank 1 of the architectural VGPR file 1108. The elements of these matrices are coupled through source multiplexers 1112 and then through architectural register rotation crossbar 1116. The outputs of architectural register rotation crossbar 1116 are coupled to double buffers 1120 for the A matrix 1002 and double buffers 1122 for the B matrix 1004. The outputs of double buffers 1120 are coupled through replica crossbar 1126 and then coupled to DOT8 cells 1130A-H. The outputs of double buffers 1122 are also coupled to DOT8 cells 1130A-H.

[0075] In one embodiment, the DOT8 units 1130A-H are configured to perform dot product or outer product operations between the rows of the A matrix 1002 and the columns of the B matrix 1004. The results of these dot product or outer product operations are coupled back to the architectural VGPR file 1108 via the multiplexer 1106. The results of the previous dot product or outer product operation marked as the source C operand in the source multiplexer 1112 can be coupled back to the input of the DOT8 units 1130A-H for further accumulation. In addition, after the A matrix 1002 and the B matrix 1004 are read from the architectural VGPR file 1108 in the first cycle, the operands can be read from the architectural VGPR file 1108 and provided to the FMA pipeline 1124 in a subsequent cycle. This allows overlapping execution on the DOT8 units 1130A-H and the FMA pipeline 1124. Note that the DOT8 units 1130A-H can also be referred to as a matrix multiplication pipeline. Furthermore, banks 2 and 3 of the architectural VGPR file 1108 may be written to the vector I / O export unit 1118 to export results generated by the DOT8 units 1130A-H or the FMA pipeline 1124 .

[0076] Now go to Fig.12 , shows one embodiment of a timing diagram 1200A for performing a matrix multiplication operation. The timing diagram 1200A illustrates a method that can be implemented to perform a matrix multiplication operation in the stream processor 1100 ( Fig.11 ) is a sequence of steps to perform a matrix multiplication operation on a processor. In cycle 0, operands of source A, source B, and source C are read from the architectural VGPR stack. In cycle 1, operands of source A, source B, and source C are provided to the matrix multiplication pipeline. Also in cycle 1, source A and source B operands may be read from the architectural VGPR stack and provided to the FMA pipeline in cycle 2 to execute a two-operand instruction. In cycle 1, the FMA pipeline is idle, but the FMA pipeline may start operations starting from cycle 2. In addition, the source D operand may be read from the accumulation VGPR stack in cycle 1 and written to memory in cycle 2. This mode of operation may continue in cycles 2-4 until the matrix multiplication pipeline completes the matrix multiplication operation. A new matrix multiplication operation may be started in cycle 5, while the FMA pipeline is idle in the 5th cycle.

[0077] Reference now Fig.13 , shows a timing diagram 1200B of another embodiment for performing a matrix multiplication operation. The timing diagram 1200B is intended to be a follow-up to the timing diagram 1200A ( Fig.12) is a continuation of the operation shown in ). In cycle 6, the matrix multiplication pipeline performs the second stage of the matrix multiplication operation, and the FMA pipeline can start a new FMA operation. Moreover, the result can be written to memory in cycle 6. In cycles 7-8, the next stage of the matrix multiplication operation can be performed by the matrix multiplication pipeline, while the FMA pipeline accesses the accumulation VGPR file of operands and performs operations overlapping with the matrix multiplication operation. This mode of operation can continue for any number of additional cycles through the matrix multiplication pipeline and the FMA pipeline.

[0078] Reference now Fig.14 , an embodiment of a method 1400 for performing a matrix multiplication operation is shown. For the purpose of discussion, the steps in this embodiment are shown in order. However, it is noted that in various embodiments of the described method, one or more elements described are performed simultaneously, performed in a different order than shown, or omitted completely. Other additional elements are also performed as needed. Any of the various systems or devices described herein is configured to implement method 1400.

[0079] The stream processor reads the first and second matrices from the first vector register file and stores the first and second matrices in a temporary memory (block 1405). Note that the first and second matrices read and stored in block 1405 may actually be part of a larger matrix. Next, the stream processor provides the first portion of the first matrix and the first portion of the second matrix to a matrix multiplication pipeline (block 1410). The matrix multiplication pipeline then generates a result that is a dot product or outer product of the elements of the first portion of the first matrix with the corresponding elements of the first portion of the second matrix (block 1415). Next, the matrix multiplication pipeline writes the result of the dot product or outer product operation to the second vector register file (block 1420).

[0080] Then, if the matrix multiplication operation is complete (decision box 1425, "yes" branch), the stream processor writes the result of the matrix multiplication operation to the memory (box 1430). After box 1430, method 1400 ends. If the matrix multiplication operation is not complete (decision box 1425, "no" branch), the stream processor provides the next part of the first matrix and the next part of the second matrix to the matrix multiplication pipeline (box 1435). The stream processor also provides the accumulated value from the second vector register file to the matrix multiplication pipeline (box 1440). In another embodiment, the accumulated value can be read from the memory and provided to the matrix multiplication pipeline. In one embodiment, the accumulated value is the result of a previous dot product operation performed by the matrix multiplication pipeline.

[0081] Next, the matrix multiplication pipeline generates a result, which is the dot product or outer product of the elements of the first matrix and the corresponding elements of the second matrix (block 1445). Moreover, the matrix multiplication pipeline adds the accumulated value to the result of the current dot product or outer product operation (block 1450). In another embodiment, the result of the current dot product or outer product operation is added to the accumulated value. The matrix multiplication pipeline then writes the sum (calculated in block 1450) to the second vector register file (block 1455). After block 1455, method 1400 returns to decision block 1425.

[0082] In various embodiments, the method and / or mechanism described herein are implemented using program instructions of a software application. For example, it is conceivable that program instructions that can be executed by a general or special processor can be implemented. In various embodiments, such program instructions can be represented by a high-level programming language. In other embodiments, the program instructions can be compiled from a high-level programming language into a binary, intermediate or other form. Alternatively, program instructions describing hardware behavior or design can be written. Such program instructions can be represented by a high-level programming language such as C. Alternatively, a hardware design language (HDL) such as Verilog can be used. In various embodiments, program instructions are stored on any one of various non-temporary computer-readable storage media. The storage medium can be accessed by a computing system during use to provide program instructions to the computing system for program execution. Generally speaking, such a computing system includes at least one or more memories and one or more processors configured to execute program instructions.

[0083] It should be emphasized that the above embodiments are only non-limiting examples of implementation schemes. Once the above disclosure is fully understood, many changes and modifications will become apparent to those skilled in the art. It is intended that the following claims be interpreted as including all such changes and modifications.

Claims

1. A system for performing a matrix multiplication operation, the system comprising: first vector register file; and a first execution pipeline coupled to the first vector register file, wherein the first execution pipeline includes a plurality of dot product units; Wherein, in order to perform a matrix multiplication operation on a first matrix including two or more dimensions and a second matrix including two or more dimensions, the system is configured as follows: fetching each column of the first matrix and each row of the second matrix into corresponding vector registers only once; using the plurality of dot product units, performing an outer product operation using values ​​stored in the first vector register and the second vector register; In a first cycle, a first portion of an intermediate matrix is ​​calculated by using a first portion of a given column of the first matrix and an entirety of a given row of the second matrix, the intermediate matrix being an outer product of the given column of the first matrix and the given row of the second matrix; in a second cycle after the first cycle, calculating a second portion of the intermediate matrix different from the first portion by using a second portion different from the first portion of the given column and the entirety of the given row; as well as The values ​​generated by the plurality of dot product units are accumulated.

2. The system of claim 1, wherein the system further comprises a second vector register file, wherein the system is configured to read a plurality of accumulated inputs from the second vector register file and provide the plurality of accumulated inputs to the first execution pipeline.

3. The system of claim 2, wherein: Each dot product unit is further configured to write an output to the second vector register file, wherein the output of a previous dot product operation is an accumulated input added to the sum of a current dot product operation.

4. The system of claim 1, wherein the first execution pipeline is further configured to: In a first cycle, fetching a given column and a given row from the first vector register file to a plurality of storage elements; and The given column and the given row are stored in the plurality of storage elements for repeated use by the plurality of dot product units in a plurality of cycles after the first cycle until each element of the intermediate matrix is ​​calculated.

5. The system of claim 4 , wherein the system further comprises a second execution pipeline, wherein the system is further configured to read elements of a third matrix from the first vector register file and provide the elements of the third matrix to the second execution pipeline in any cycle of the plurality of cycles after the first cycle.

6. The system of claim 1, wherein the first execution pipeline further comprises a crossbar switch configured to rotate the elements of the first matrix before sending the elements of the first matrix to the plurality of dot product units, while the elements of the second matrix remain unchanged for the plurality of dot product units.

7. A method for performing a matrix multiplication operation, comprising: Performing a matrix multiplication operation on a first matrix including two or more dimensions and a second matrix including two or more dimensions by a first execution pipeline of the plurality of execution pipelines in the following manner: fetching each column of the first matrix and each row of the second matrix into corresponding vector registers only once; Using a plurality of dot product units, performing an outer product operation using values ​​stored in the first vector register and the second vector register; In a first cycle, a first portion of an intermediate matrix is ​​calculated by using a first portion of a given column of the first matrix and an entirety of a given row of the second matrix, the intermediate matrix being an outer product of the given column of the first matrix and the given row of the second matrix; in a second cycle after the first cycle, calculating a second portion of the intermediate matrix different from the first portion by using a second portion different from the first portion of the given column and the entirety of the given row; as well as The values ​​generated by the plurality of dot product units are accumulated.

8. The method of claim 7, further comprising reading a plurality of accumulated inputs from a second vector register file and providing the plurality of accumulated inputs to the first execution pipeline.

9. The method of claim 8, further comprising writing output to the second vector register file, wherein the output of a previous dot product operation is an accumulated input added to the sum of a current dot product operation.

10. The method of claim 7, further comprising: In a first cycle, fetching a given column and a given row from a first vector register file to a plurality of storage elements; and The given column and the given row are stored in the plurality of storage elements for repeated use by the plurality of dot product units in a plurality of cycles after the first cycle until each element of the intermediate matrix is ​​calculated.

11. The method of claim 10, further comprising reading elements of a third matrix from the first vector register file and providing the elements of the third matrix to a second execution pipeline in any cycle of the plurality of cycles after the first cycle.

12. The method of claim 7, further comprising rotating the elements of the first matrix before sending the elements of the first matrix to the plurality of dot product units through a crossbar of the first execution pipeline, while the elements of the second matrix remain unchanged for the plurality of dot product units.

13. A device for performing a matrix multiplication operation, the device comprising: Multiple vector register files; and a plurality of execution pipelines coupled to the plurality of vector register files; a plurality of dot product units in a first execution pipeline of the plurality of execution pipelines; Wherein, in order to perform a matrix multiplication operation on a first matrix including two or more dimensions and a second matrix including two or more dimensions, the apparatus is configured to: fetching each column of the first matrix and each row of the second matrix into corresponding vector registers only once; using the plurality of dot product units, performing an outer product operation using values ​​stored in a first vector register and the vector register; In a first cycle, a first portion of an intermediate matrix is ​​calculated by using a first portion of a given column of the first matrix and an entirety of a given row of the second matrix, the intermediate matrix being an outer product of the given column of the first matrix and the given row of the second matrix; in a second cycle after the first cycle, calculating a second portion of the intermediate matrix different from the first portion by using a second portion different from the first portion of the given column and the entirety of the given row; as well as The values ​​generated by the plurality of dot product units are accumulated.

14. The apparatus of claim 13, wherein the apparatus is configured to read a plurality of accumulated inputs from a second vector register file and provide the plurality of accumulated inputs to the first execution pipeline.

15. The device of claim 14, wherein: The apparatus is configured to write an output to the second vector register file, wherein an output of a previous dot product operation is an accumulated input added to a sum of a current dot product operation.

16. The device of claim 13, wherein: The device is also configured to: In a first cycle, fetching a given column and a given row from a first vector register file to a plurality of storage elements; and The given column and the given row are stored in the plurality of storage elements for repeated use by the plurality of dot product units in a plurality of cycles after the first cycle until each element of the intermediate matrix is ​​calculated.

17. The apparatus of claim 16, wherein the apparatus is further configured to read elements of a third matrix from the first vector register file and provide the elements of the third matrix to a second execution pipeline in any cycle of the plurality of cycles after the first cycle.

Citation Information

Patent Citations

  • Multiplying and adding matrices

    US20110153707A1