Tiling algorithm for matrix math instruction sets

By converting matrix elements from linear format to tiled format and using multiple computing units and caches for parallel processing, the problem of low matrix computing efficiency in parallel processing units is solved, and more efficient memory bandwidth utilization and matrix computing performance are achieved.

CN111338974BActive Publication Date: 2025-05-16ADVANCED MICRO DEVICES INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201811566422.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2018-12-19
Publication Date
2025-05-16
Estimated Expiration
2038-12-19

AI Technical Summary

Technical Problem

When performing matrix operations in parallel processing units, the prior art requires complex formulas to calculate offsets, resulting in inefficient use of memory bandwidth and matrix computing units.

Method used

Reliance on memory bandwidth is reduced by converting matrix elements stored in memory from linear format to tiled format and processing in parallel using multiple computing units and caches.

Benefits of technology

This improves the efficiency of matrix operations, reduces the complexity of calculating offsets, and improves the utilization of memory bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0001911861630000011
    Figure HDA0001911861630000011
  • Figure HDA0001911861630000021
    Figure HDA0001911861630000021
  • Figure HDA0001911861630000031
    Figure HDA0001911861630000031
Patent Text Reader

Abstract

The present invention relates to a tiled algorithm for a matrix math instruction set. A system, an apparatus and a method for implementing a tiled algorithm for a matrix math instruction set are disclosed. A system includes at least a memory, a cache, a processor and a plurality of computing units. The memory stores a plurality of matrix elements in a linear format, and the processor converts the plurality of matrix elements from the linear format to a tiled format. Each computing unit retrieves a plurality of matrix elements from the memory into a cache. Each computing unit includes a matrix operation unit that loads a plurality of matrix elements of a corresponding tile from the cache and performs a matrix operation on the plurality of matrix elements to generate a result in a tiled format. The system implements classification of a first data set based on the result of the matrix operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates generally to the field of computers, and more particularly to a tiling algorithm for a matrix math instruction set. Background Art

[0002] Performing matrix operations in a parallel processing unit involves loading a large amount of data from memory, which results in high memory bandwidth utilization. Loading matrix elements usually requires calculating offsets in order to step through the matrix data elements stored in a linear format in memory. However, this requires complex formulas for each load instruction to calculate the offsets for moving through the elements of the matrix in the correct order to perform matrix operations. As used herein, the term "linear format" is defined as a format in which continuous matrix elements are stored in adjacent memory locations in a sequential manner until the end of a physical row or column of the memory is reached. Examples of linear formats for storing matrix elements include row-major order and column-major order. In row-major order, continuous elements of a matrix row are adjacent to each other in memory. In column-major order, continuous elements of a matrix column are adjacent to each other in memory. Processing matrix elements in a linear format results in inefficient use of memory bandwidth and matrix operation units. Summary of the invention

[0003] Some aspects of the invention can be described as follows:

[0004] 1. A system, which may include:

[0005] a memory storing a plurality of matrix elements in a linear format;

[0006] Cache;

[0007] a processor configured to convert the plurality of matrix elements from the linear format to a tiled format; and

[0008] A plurality of computing units, wherein each computing unit of the plurality of computing units comprises a matrix operation unit;

[0009] Each matrix operation unit is configured to:

[0010] loading a given plurality of matrix elements of one or more corresponding tiles from the cache; and

[0011] performing a matrix operation on the given plurality of matrix elements to generate a result in the tiled format;

[0012] The system is configured to implement classification of the first data set based on multiple results from multiple matrix operation units.

[0013] 2. The system of clause 1, wherein the cache comprises a plurality of channels, wherein the plurality of computational units are configured to load matrix elements on the plurality of channels in parallel in a single clock cycle.

[0014] 3. A system according to clause 2, wherein:

[0015] Each computing unit is further configured to execute kernels in parallel that are equivalent to kernels executed by other computing units in the plurality of computing units; and

[0016] Each matrix operation unit is configured to load matrix elements from the cache on a different channel from the other matrix operation units.

[0017] 4. The system of clause 1, wherein in response to receiving a request to convert a plurality of matrix elements of a first source matrix from the linear format to a tiled format, the processor is configured to:

[0018] reading values ​​from sequential locations of a first buffer in the memory, wherein the first buffer stores matrix elements in the linear format; and

[0019] When writing the value to a second buffer, stepping through the second buffer with a stride equal to a tile height, wherein the second buffer stores matrix elements in the tiled format.

[0020] 5. A system according to clause 1, wherein the matrix operation is a matrix multiplication operation, and wherein the given plurality of matrix elements are transmitted along a plurality of lanes of each matrix operation unit.

[0021] 6. A system according to clause 1, wherein the classification of the first data set is implemented during execution of a machine learning engine application.

[0022] 7. The system of clause 1, wherein elements of the first source matrix stored in consecutive memory locations in the linear format are stored in a tiled format in memory locations separated by a tile height.

[0023] 8. A method, which may include:

[0024] converting, by a processor, a plurality of matrix elements stored in a memory from a linear format to a tiled format;

[0025] loading a plurality of matrix elements from the memory into a cache memory via a plurality of computational units;

[0026] loading, by each of a plurality of matrix operation units, a given plurality of matrix elements of one or more corresponding tiles from the cache;

[0027] By each matrix operation unit, performing a matrix operation on the given plurality of matrix elements to generate a result in the tiled format; and

[0028] Classification of the first data set is achieved based on the plurality of results from the plurality of matrix operation units.

[0029] 9. The method of clause 8, wherein the cache comprises a plurality of channels, and wherein the method further comprises loading matrix elements on the plurality of channels in parallel in a single clock cycle by the plurality of computational units.

[0030] 10. The method according to clause 9, further comprising:

[0031] executing, in parallel, by each computing unit, a kernel that is equivalent to kernels executed by other computing units of the plurality of computing units; and

[0032] Matrix elements from the cache are loaded by each matrix operation unit on a different channel from the other matrix operation units.

[0033] 11. A method according to clause 8, wherein in response to receiving a request to convert a plurality of matrix elements of a first source matrix from the linear format to a tiled format, the method comprises:

[0034] reading values ​​from sequential locations of a first buffer in the memory, wherein the first buffer stores matrix elements in the linear format; and

[0035] When writing the value to a second buffer, stepping through the second buffer with a stride equal to a tile height, wherein the second buffer stores matrix elements in the tiled format.

[0036] 12. The method according to clause 8 further comprises: for each matrix operation unit, transmitting the given plurality of matrix elements along a plurality of lanes of the matrix operation unit, wherein the matrix operation is a matrix multiplication operation.

[0037] 13. A method according to clause 8, wherein the classification of the first data set is implemented during execution of a machine learning engine application.

[0038] 14. The method of clause 8, wherein elements of the first source matrix stored in the linear format in consecutive memory locations are stored in a tiled format in memory locations separated by a tile height.

[0039] 15. An apparatus, which may include:

[0040] a memory storing a plurality of matrix elements in a linear format; and

[0041] A plurality of computing units, the plurality of computing units being configured to:

[0042] generating a request to convert the plurality of matrix elements from the linear format to a tiled format;

[0043] performing a matrix operation on the plurality of matrix elements in the tiled format to generate a plurality of results in the tiled format; and

[0044] Classification of the first data set is achieved based on the plurality of results.

[0045] 16. The apparatus of clause 15, further comprising a cache, wherein the cache comprises a plurality of channels, and wherein the plurality of computational units are configured to load matrix elements on the plurality of channels in parallel in a single clock cycle.

[0046] 17. The apparatus of clause 16, wherein each of the plurality of computing units is configured to:

[0047] executing kernels in parallel that are equivalent to kernels executed by other computing units of the plurality of computing units; and

[0048] Matrix elements are loaded from the cache onto a channel different from other channels utilized by other computational units of the plurality of computational units.

[0049] 18. An apparatus according to clause 15, wherein in response to receiving a request to convert a plurality of matrix elements of a first source matrix from the linear format to a tiled format, the apparatus is configured to:

[0050] reading values ​​from sequential locations of a first buffer in the memory, wherein the first buffer stores matrix elements in the linear format; and

[0051] When writing the value to a second buffer, stepping through the second buffer with a stride equal to a tile height, wherein the second buffer stores matrix elements in the tiled format.

[0052] 19. The apparatus of clause 15, wherein matrix elements are transmitted along a plurality of lanes of each of the plurality of computational units, and wherein the matrix operation is a matrix multiplication operation.

[0053] 20. An apparatus according to clause 15, wherein the classification of the first data set is implemented during execution of a machine learning engine application. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Advantages of the methods and mechanisms described herein may be better understood by referring to the following description in conjunction with the accompanying drawings, in which:

[0055] Figure 1 It is a block diagram of one implementation of a computing system.

[0056] Figure 2 is a block diagram of another implementation of a computing system.

[0057] Figure 3 It is a block diagram of one implementation of a matrix operation unit.

[0058] Figure 4 Schematic diagram of one implementation of the data layout of the source A matrix operated by the SIMD unit.

[0059] Figure 5 Schematic diagram of one implementation of the data layout of the source B matrix operated by the SIMD unit.

[0060] Figure 6 is a schematic diagram of one implementation of a tiled layout for tiled blocks.

[0061] Figure 7 It is a schematic diagram of one implementation method of the format of 32×32 blocks in the source A matrix.

[0062] Figure 8 It is a schematic diagram of one implementation of the format of 128×128 blocks within the source A matrix.

[0063] Fig. 9 It is a schematic diagram of one implementation of a format of an image consisting of 128×128 blocks.

[0064] Fig.10 It is a schematic diagram of one implementation of the source B matrix.

[0065] Fig.11 It is a schematic diagram of one implementation of the format of a 32×32 block within a source B matrix.

[0066] Fig.12 It is a schematic diagram of one implementation of the format of 128×128 blocks within the source B matrix.

[0067] Fig.13 It is a schematic diagram of one implementation of a format of an image consisting of 128×128 blocks.

[0068] Fig.14 is a schematic diagram of one implementation of the resulting C matrix.

[0069] Fig.15An example of pseudo code for converting a source A matrix from a linear format to a tiled format is shown according to one implementation.

[0070] Fig.16 An example of pseudo code for converting a source B matrix from a linear format to a tiled format is shown according to one implementation.

[0071] Fig.17 is a generalized flow chart illustrating one implementation of a method for converting matrix data from a linear format to a tiled format to perform matrix operations. DETAILED DESCRIPTION

[0072] In the following description, many specific details are set forth to provide a thorough understanding of the methods and mechanisms presented herein. However, those of ordinary skill in the art will recognize that various implementations can be practiced without these specific details. In some cases, well-known structures, components, signals, computer program instructions, and techniques are not shown in detail to avoid making the methods described herein difficult to understand. It should be understood that, for simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the size of some elements may be enlarged relative to other elements.

[0073] Various systems, devices and methods for converting matrix data from a linear format to a tiled format to perform matrix operations are disclosed herein. A system includes at least a memory, a cache, a processor and a plurality of computing units. The memory stores a plurality of matrix elements in a linear format, and the processor converts the plurality of matrix elements from a linear format to a tiled format. In one implementation, the term "tile format" or "tile format" is defined as a format for storing matrix elements in a memory location so that the matrix elements used to perform a given matrix operation are stored in adjacent memory locations. In one implementation, the given matrix operation is a matrix multiplication operation. In another implementation, the term "tile format" is defined as a format in which the matrix elements constituting the columns of a tile are stored in adjacent memory locations. The term "tile" is defined as an N×M block of elements, where N and M are positive integers, and where at least one of N or M is greater than 1. "Tile" may also be referred to herein as "block". In one implementation, the number of columns "M" of the tile of the first source matrix is ​​equal to the number of channels of the matrix operation unit. In one implementation, for the matrix multiplication operation, the first source matrix is ​​divided into tiles of N×M elements, and the second source matrix is ​​divided into tiles of M×N elements.

[0074] The tiled format causes the matrix data to be ordered in a specific layout that allows data to be loaded continuously without performing offset calculations. Each computing unit retrieves rows and / or columns of matrix elements of one or more tiles from memory to enter them into a cache. Each computing unit includes a matrix operation unit that retrieves multiple matrix elements of the corresponding tile from the cache and performs matrix operations on the multiple matrix elements to generate a result in the tiled format. In one implementation, the system performs multiple matrix operations on multiple computing units that are part of a machine learning engine to achieve classification of a first data set. For example, in one implementation, the system performs multiple matrix operations while implementing a neural network to classify an image into one or more categories. The neural network can be a convolutional neural network, a recursive neural network, or other types. Various tasks such as handwritten digit classification and face detection can be performed by the neural network. In addition, the neural network can perform other more challenging visual classification tasks. Other applications of neural networks include speech recognition, language modeling, sentiment analysis, text prediction, etc. In other implementations, the system performs multiple matrix operations on multiple computing units that are part of other types of software applications.

[0075] In various implementations, the cache has P channels, where P is a positive integer greater than 1. In one implementation, P is equal to 32. In other implementations, P is equal to other numbers. Requests are mapped to different channels of the cache based on a portion of the physical address bits. For example, in one implementation, each channel is mapped with bits 12-8 of the physical address. In other implementations, channel mapping is based on other bits of the physical address. In one implementation, storing matrix elements in a tiled format increases cache hit efficiency. In a typical application, each computing unit processes a different matrix tile, but the tile can be mapped to the same cache channel. This affects cache efficiency because different computing units will eventually request data through the same cache channel. Therefore, the computing unit will wait for data to be returned from the cache, and the cache will process requests one by one in the same channel. However, when the matrix elements are stored in a tiled format, different computing units are mapped to different channels. When the computing units execute the same kernel in parallel, requests will be sent to the cache through different channels, which helps to improve the access efficiency of the cache.

[0076] Reference now Figure 1, a block diagram of one implementation of a computing system 100 is shown. In one implementation, computing system 100 includes at least processors 105A-N, input / output (I / O) interface 120, bus 125, memory controller 130, network interface 135, memory device 140, display controller 150, and display 155. In other implementations, computing system 100 includes other components and / or computing system 100 is arranged in a different manner. Processors 105A-N represent any number of processors included in system 100.

[0077] In one implementation, processor 105A is a general purpose processor, such as a central processing unit (CPU). In one implementation, processor 105N is a data parallel processor with a highly parallel architecture. Data parallel processors include graphics processing units (GPUs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc. In some implementations, processors 105A-N include multiple data parallel processors. In one implementation, processor 105N is a GPU that provides pixels to display controller 150 to be driven to display 155.

[0078] The memory controller 130 represents any number and type of memory controllers accessible by the processors 105A-N. The memory controller 130 is coupled to any number and type of memory devices 140. The memory devices 140 represent any number and type of memory devices. For example, the memory types in the memory devices 140 include dynamic random access memory (DRAM), static random access memory (SRAM), NAND flash memory, NOR flash memory, ferroelectric random access memory (FeRAM), etc.

[0079] The I / O interface 120 represents any number and type of I / O interfaces (e.g., a peripheral component interconnect (PCI) bus, a PCI-Extension (PCI-X), a PCIE (PCI Express) bus, a Gigabit Ethernet (GBE) bus, a Universal Serial Bus (USB)). Various types of peripheral devices (not shown) are coupled to the I / O interface 120. These peripheral devices include (but are not limited to) a display, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, etc. The network interface 135 is used to receive and send network messages over a network.

[0080] In various implementations, computing system 100 is a computer, a laptop, a mobile device, a gaming console, a server, a streaming device, a wearable device, or any of a variety of other types of computing systems or devices. Note that the number of components of computing system 100 varies from implementation to implementation. For example, in other implementations, the number of each component is less than 100. Figure 1 It should also be noted that in other implementations, the computing system 100 includes Figure 1 Other components not shown in FIG. In addition, in other implementations, the computing system 100 is used with Figure 1 The method shown in the figure is different from other methods of construction.

[0081] Now turn to Figure 2 , a block diagram of another implementation of a computing system 200 is shown. In one implementation, system 200 includes a GPU 205, a system memory 225, and a local memory 230. System 200 also includes other components not shown to avoid obscuring the figure. GPU 205 includes at least a command processor 235, a control logic 240, a scheduling unit 250, a computing unit 255A-N, a memory controller 220, a global data sharing 270, a level 1 (L1) cache 265, and a level 2 (L2) cache 260. In other embodiments, GPU 205 includes other components, omits one or more of the components shown, has multiple instances of components, although Figure 2 Only one example is shown in FIG. 1 , and / or the circuits of GPU 205 may be organized in other suitable ways. In one implementation, the circuits of GPU 205 include ( Figure 1 ) processor 105N.

[0082] In various implementations, the computing system 200 executes any of a variety of types of software applications. As part of executing a given software application, a host CPU (not shown) of the computing system 200 launches a kernel to be executed on the GPU 205. The command processor 235 receives the kernel from the host CPU and uses the scheduling unit 250 to issue a corresponding wavefront to the compute units 255A-N. In one implementation, each compute unit 255A-N includes a matrix operation unit. For example, the matrix operation unit performs matrix multiplication operations. Additionally, in various implementations, the matrix operation unit performs other types of matrix operations. The wavefronts executed on the compute units 255A-N read and write data to the global data share 270, L1 cache 265, and L2 cache 260 within the GPU 205. Although not shown in Figure 2, but in one implementation, the computing units 255A-N also include one or more caches and / or local memories within each computing unit 255A-N.

[0083] In one implementation, the matrix data is stored in a linear format in the system memory 225 and / or the local memory 230. Before loading the matrix data into the L2 cache 260, the matrix data is converted from the linear format to a tiled format. In one implementation, the term "tile format" is defined as storing matrix elements together into cells of tiles, with each tile storing N×M blocks of matrix elements, where N and M are positive integers. The "tile format" causes consecutive tiles to be stored in memory in a sequential manner.

[0084] In one implementation, the command processor 235 converts the matrix data from a linear format to a tiled format. In another implementation, a host processor (e.g., Figure 1 The processor 105A of the processor 105A) converts the matrix data from a linear format to a tiled format. Then, during the execution of a wavefront on the computing unit 255A-N, the matrix data is loaded from the L2 cache 260 in an efficient manner because the matrix data is stored in a tiled format. In one implementation, when the matrix elements are stored in a tiled format, the matrix data elements are read into the computing unit 255A-N in parallel from the L2 cache 260 on multiple channels. This is a more efficient method than storing the matrix data elements in a linear format, which would cause the matrix data elements to be read from the L2 cache 260 on a single channel in a serial manner.

[0085] Reference now Figure 3 , a block diagram of one implementation of a matrix operation unit 300 is shown. In one implementation, each computation unit 255A-N includes circuitry of a matrix operation unit 300. In one implementation, the matrix operation unit 300 includes at least an architecture vector general purpose register (VGPR) file 305, a derivation unit 310, an accumulation VGPR file 315, a single instruction, multiple data (SIMD) unit 320, and a SIMD unit 325. It should be understood that the matrix operation unit 300 includes any number of other components that are not shown in order to avoid obscuring the figure. Additionally, in other implementations, the matrix operation unit 300 is organized in other suitable ways.

[0086] In one implementation, SIMD unit 320 is a floating point unit for performing various floating point operations, and SIMD unit 325 is a matrix unit for performing various matrix operations (e.g., dot product operations, matrix multiplication operations). In one implementation, each individual input connected to the architecture VGPR file 305 and the accumulation VGPR file 315 is shown to have 16 channels, each channel has 32 bits. In other implementations, the input has other numbers of channels with other bit widths. In one implementation, when the matrix elements are stored in a tiled format, the SIMD unit 325 operates on the input matrix elements more efficiently. Therefore, in this implementation, the matrix elements are converted from a linear format to a tiled format before being read into the architecture VGPR file 305 and / or the accumulation VGPR file 315. This enables the operation to be performed by the SIMD unit 325 in a more efficient manner.

[0087] Now turn to Figure 4 , shows a schematic diagram of an implementation of a data layout 400 of a source A matrix operated by a SIMD unit. In one implementation, the SIMD unit 325 ( Figure 3 )according to Figure 4 400 is organized for reading in the source A matrix to perform a matrix multiplication operation of the source A matrix multiplied by the source B matrix. For example, in this implementation, each SIMD unit has 64 threads in the data layout 400. In other implementations, the SIMD unit includes other numbers (e.g., 32, 128) of threads in the data layout. Each thread in the data layout 400 corresponds to a channel of the SIMD unit.

[0088] Depending on the size of the blocks being processed, a different number of blocks may be mapped according to data layout 400. For example, if the size of each block is 32×2, two blocks (Blk0 and Blk1) are mapped from the VGPR file to the SIMD unit. If the size of each block is 16×2, four blocks (Blk0, Blk1, Blk2, and Blk3) are mapped from the VGPR file to the SIMD unit. If the size of each block is 4×2, 16 blocks (Blk0, Blk1, Blk2, Blk3, Blk4, Blk5, Blk6, Blk7, Blk8, Blk9, Blk10, Blk11, Blk12, Blk13, Blk14, and Blk15) are mapped from the VGPR file to the SIMD unit.

[0089] Reference now Figure 5 , a schematic diagram of an implementation of a data layout 500 for a source B matrix of a SIMD unit is shown. In one implementation, a VGPR file and a SIMD unit (e.g., SIMD unit 325 ( Figure 3 )). For example, in one implementation, data layout 500 defines connections for loading source B matrices so as to perform a matrix multiplication operation between source A matrices and source B matrices. In one implementation, data layout 500 is organized for 64 threads. In other implementations, data layout 500 includes other numbers of threads. Each thread of data layout 500 corresponds to a lane of a SIMD unit. Depending on the size of the block being processed (e.g., 2×32, 2×16, 2×4), different numbers of blocks (e.g., 2, 4, 16) can be mapped according to data layout 500.

[0090] Now turn to Figure 6 , a schematic diagram of an implementation of a tiled layout of a block 600 of a source A matrix is ​​shown. In one implementation, the matrix is ​​divided into a plurality of blocks, each of which is organized according to a block 600. The organization of the elements in the block 600 shows the mapping relationship between linear data and tiled data. For example, the matrix element 65 is represented by a circle 605. The original position of the element is (x=3, y=0). Therefore, the position in the linear format is (y*stride+x)=0*4+3=3, and the position in the tiled format is 65. In one implementation, the basic tiled block is 32×4 for the source A matrix and 4×32 for the source B matrix to implement matrix multiplication operations. In other implementations, the basic tiled block of the matrix operation can be of other sizes. In one implementation, for the next level in the matrix hierarchy, 8 basic blocks of size 32×4 are combined into an intermediate block of 32×32. In addition, moving up the matrix hierarchy, 16 intermediate blocks are combined into a large block of 128×128. For the purposes of the remaining discussion, it is assumed that the matrix size being processed is aligned with a large block of 128 x 128. However, it should be understood that in other implementations, the matrix size may be aligned with blocks of other sizes.

[0091] Reference now Figure 7 , a diagram of one implementation of a 32×32 block 700 within a source A matrix is ​​shown. As shown, each 4×32 block is arranged from left to right from block 0 to block 7 to form a 32×32 block 700. In other implementations, blocks of other sizes may be combined together to form blocks higher in the matrix block hierarchy.

[0092] Now turn to Figure 8, a schematic diagram of one implementation of a 128×128 block 800 within a source A matrix is ​​shown. The first column of the 128×128 block 800 is composed of 32×32 blocks 0-3, as shown in block 800. The second column of the 128×128 block 800 is composed of 32×32 blocks 4-7, the third column of the 128×128 block 800 is composed of 32×32 blocks 8-12, and the fourth column (i.e., the rightmost column) of the 128×128 block 800 is composed of 32×32 blocks 12-15. In one implementation, according to block 700 ( Figure 7 ) construct each 32×32 block.

[0093] Reference now Fig. 9 , shows a schematic diagram of one implementation of an image 900 composed of 128×128 blocks. It should be understood that in other implementations, the image can be composed of blocks of other sizes, and the blocks can be organized in other suitable ways. For example, in other implementations, the image can be composed of 64×64 blocks, 256×256 blocks, or blocks of other sizes. For these implementations, the mechanisms and methods described here for image 900 can be adjusted to apply to other images composed of blocks of other sizes.

[0094] Now turn to Fig.10 , a schematic diagram of one implementation of a tiled layout of a source B matrix block 1000 is shown. In one implementation, the matrix operation unit multiplies the source A matrix by the source B matrix. The organization of the source B matrix block 1000 is one example of the organization of the elements of the source B matrix block 1000 according to one implementation for implementing a matrix multiplication operation.

[0095] Reference now Fig.11 , shows a schematic diagram of an implementation of a 32×32 block 1100 of a source B matrix. In one implementation, eight different 32×4 blocks are used to construct the 32×32 block 1100 of the source B matrix. Fig.11 As shown, each 32×4 block is arranged from top to bottom from block 0 to block 7, as shown in block 1100. In other implementations, 32×32 block 1100 may be constructed with other numbers, sizes, and / or arrangements of smaller blocks.

[0096] Now go to Fig.12, a schematic diagram of one implementation of a 128×128 block 1200 of a source B matrix is ​​shown. The first row of the 128×128 block 1200 is composed of 32×32 blocks 0-3 moving from left to right. The second row of the 128×128 block 1200 is composed of 32×32 blocks 4-7 moving from left to right, the third row of the 128×128 block 1200 is composed of 32×32 blocks 8-11 moving from left to right, and the fourth row (i.e., the bottom row) of the 28×128 block 1200 is composed of 32×32 blocks 12-15 moving from left to right. In one implementation, according to the block 1100 ( Fig.11 ) layout to construct each 32×32 block.

[0097] Reference now Fig.13 , a schematic diagram of one implementation of an image 1300 composed of 128×128 blocks is shown. The first row of image 1300 includes 128×128 blocks 0-7 moving from left to right. Similarly, the second row of image 1300 includes 128×128 blocks 8-15 moving from left to right. The other rows of image 1300 are organized in the same manner, with the bottom row including 128×128 blocks 56-63 moving from left to right. In one implementation, according to block 1200 ( Fig.12 ) construct each 128×128 block 0-63.

[0098] Now turn to Fig.14 , a schematic diagram of an implementation of a block 1400 of a result C matrix is ​​shown. In one implementation, a matrix operation unit multiplies a source A matrix by a source B matrix to generate a result C matrix. The organization of the elements of block 1400 is an example of organizing the elements of a block within a result C matrix after a matrix multiplication operation has been performed according to one implementation. The first column of block 1400 includes element 0, followed by element 1, element 2, and then element 3. The second column of block 1400 includes element 4, element 5, element 6, and element 7, wherein element 4 is at the top of element 5, element 5 is at the top of element 6, and element 6 is at the top of element 7. This pattern of element layout continues in columns moving to the right of matrix C 1400 until the last rightmost column includes element 60, element 61, element 62, and element 63.

[0099] Reference now Fig.15 , shows an example of pseudo code 1500 for converting a source A matrix from a linear format to a tiled format according to one implementation. Pseudo code 1500 includes definitions of variables with specific values ​​for one specific implementation. In other implementations, the values ​​of these variables may vary depending on the size of the tile, the number of channels and bit width of the matrix multiplication unit, the number of cache channels, etc.

[0100] To discuss pseudo code 1500 and pseudo code 1600 ( Fig.16 ), assume that there are two input matrices: a source A matrix and a source B matrix. Also assume that the size of the source A matrix is ​​M×K, the size of the source B matrix is ​​K×N, and the size of the result C matrix is ​​M×N, where M, N, and K are positive integers. In one implementation, M, N, and K are equal to 1024. In other implementations, the values ​​of M, N, and K may vary.

[0101] In one implementation, there are two groups of buffers in the memory. A_outbuffer[] and B_outbuffer[] store the matrix elements of the source A matrix and the source B matrix in the memory in a linear format, respectively. A_package_outbuffer[] and B_package_outbuffer[] store the matrix elements of the source A matrix and the source B matrix in the memory in a tiled format, respectively. Based on the instructions in pseudocode 1500, the elements of the source A matrix stored in continuous memory locations in a linear format are stored in a tiled format in memory locations that separate the tile height. In other words, values ​​are read from continuous locations of the first buffer (i.e., A_outbuffer[]) and written to locations in the second buffer (i.e., A_package_outbuffer[]), while stepping through the second buffer with a stride equal to the tile height. After the source A matrix and the source B matrix are converted from a linear format to a tiled format, the kernel code executed on the computing unit loads the data from the memory into a cache (e.g., an L2 cache).

[0102] Now turn to Fig.16 , shows an example of pseudo code 1600 for converting a source B matrix from a linear format to a tiled format. The discussion of pseudo code 1600 is intended to be used as a reference to pseudo code 1500 ( Fig.15 ) is a continuation of the discussion of FIG. 16A . In one implementation, pseudo code 1600 is used to convert a source B matrix from a linear format to a tiled format. Once the source B matrix is ​​converted from a linear format to a tiled format, the kernel code loads the data from the memory into the cache.

[0103] Reference now Fig.17 , an implementation of a method 1700 for converting matrix data from a linear format to a tiled format to perform matrix operations is shown. For the purpose of discussion, the steps in the implementation are shown in order. However, it should be noted that in various implementations of the described method, one or more of the described elements are performed simultaneously, in a different order than shown, or are omitted entirely. Other additional elements are also performed as needed. Any of the various systems or devices described herein is configured to implement method 1700.

[0104] A host processor (e.g., a CPU) detects a request to perform a matrix operation on matrix data stored in a linear format (block 1705). In one implementation, the matrix operation is a matrix multiplication operation. In other implementations, other types of matrix operations are requested. Next, in response to detecting the request, the host processor converts the matrix data from a linear format to a tiled format (block 1710). Examples of pseudocode 1500 and 1600 for converting a source A matrix and a source B matrix from a linear format to a tiled format are provided in FIG. Fig.15 and Fig.16 In other implementations, other techniques for converting the source A matrix and the source B matrix from a linear format to a tiled format may be used.

[0105] Next, a plurality of matrix operation units load the matrix data in a tiled format from the memory into all N channels of the cache (block 1715). In one implementation, the cache is an L2 cache. In one implementation, the cache has N channels, where N is a positive integer greater than 1. Then, the plurality of matrix operation units perform matrix operations on the matrix data in parallel to produce a result (block 1720). Next, the processor uses the result to complete a first action associated with a given software application (block 1725). In one implementation, the first action is a classification of a first data set, and the given software application is a machine learning application. In one implementation, the first data set is an image, and the classification identifies a given category to which the image belongs. In another implementation, the first data set is a video, and the classification assigns the video to a given category. In other implementations, the first data set includes other types of data classified in other ways. In other implementations, other types of actions associated with other types of software applications are performed. After block 1725, method 1700 ends.

[0106] In various implementations, the method and / or mechanism described herein are implemented using program instructions of a software application. For example, it is conceivable that program instructions that can be executed by a general or special processor are executed. In various implementations, such program instructions are represented by a high-level programming language. In other implementations, the program instructions are compiled from a high-level programming language into a binary, intermediate or other form. Optionally, program instructions describing hardware behavior or design are written. Such program instructions can be represented by a high-level programming language such as C. Optionally, a hardware design language (HDL) such as Verilog is used. In various implementations, program instructions are stored on any one of various non-temporary computer-readable storage media. The storage medium can be accessed by a computing system during use to provide program instructions to the computing system for program execution. Generally speaking, such a computing system includes at least one or more memories and one or more processors configured to execute program instructions.

[0107] It should be emphasized that the above implementations are only non-limiting examples of implementations. For those skilled in the art, many changes and modifications will become apparent once the above disclosure is fully understood. It is intended that the following claims be interpreted as including all such changes and modifications.

Claims

1. A system comprising: a memory storing a plurality of matrix elements in a linear format; a cache, the cache comprising a plurality of channels; a processor configured to convert the plurality of matrix elements from the linear format to a tiled format; as well as a plurality of computing units, wherein each computing unit of the plurality of computing units comprises a matrix operation unit, the matrix operation unit comprising one or more single instruction multiple data (SIMD) units; Wherein when two or more computing units of the plurality of computing units execute the same kernel in parallel, each of the plurality of channels of the SIMD unit of each matrix operation unit is configured to: loading in parallel a given plurality of matrix elements of one or more corresponding tiles from the cache through different ones of the plurality of channels of the cache; as well as A matrix operation is performed on the given plurality of matrix elements to generate a result in the tiled format. 2 . The system of claim 1 , wherein the plurality of computational units are configured to load matrix elements in parallel through the plurality of channels in a single clock cycle.

3. The system of claim 1 , wherein in response to receiving a request to convert a plurality of matrix elements of a first source matrix from the linear format to a tiled format, the processor is configured to: reading values ​​from sequential locations of a first buffer in the memory, wherein the first buffer stores matrix elements in the linear format; and When writing the value to a second buffer, stepping through the second buffer with a stride equal to a tile height, wherein the second buffer stores matrix elements in the tiled format.

4. The system of claim 1, wherein the matrix operation is a matrix multiplication operation, and wherein the given plurality of matrix elements are transmitted along a plurality of lanes of each matrix operation unit. 5 . The system of claim 1 , wherein the system is configured to generate a classification of the first data set based on a plurality of results from a plurality of matrix operation units. 6 . The system of claim 1 , wherein elements of the first source matrix stored in consecutive memory locations in the linear format are stored in a tiled format in memory locations separated by a tile height.

7. A method comprising: converting, by a processor, a plurality of matrix elements stored in a memory from a linear format to a tiled format; storing the plurality of matrix elements in the tiled format in a cache, the cache comprising a plurality of channels; The same kernel is executed in parallel by two or more computing units; loading in parallel a given plurality of matrix elements of one or more corresponding tiles from the cache through different ones of the plurality of channels of the cache; as well as Through each matrix operation unit, a matrix operation is performed on the given plurality of matrix elements to generate a result in the tiled format. 8 . The method of claim 7 , wherein the method further comprises loading matrix elements on the plurality of channels in parallel in a single clock cycle by the plurality of computational units.

9. The method of claim 7, wherein in response to receiving a request to convert a plurality of matrix elements of a first source matrix from the linear format to a tiled format, the method comprises: reading values ​​from sequential locations of a first buffer in the memory, wherein the first buffer stores matrix elements in the linear format; as well as When writing the value to a second buffer, stepping through the second buffer with a stride equal to a tile height, wherein the second buffer stores matrix elements in the tiled format.

10. The method according to claim 7, further comprising: For each matrix operation unit, the given plurality of matrix elements are transmitted along a plurality of channels of the matrix operation unit, wherein the matrix operation is a matrix multiplication operation.

11. The method of claim 7, further comprising generating a classification of the first data set based on a plurality of results from a plurality of matrix operation units.

12. The method of claim 7, wherein elements of the first source matrix stored in the linear format in consecutive memory locations are stored in a tiled format in memory locations separated by a tile height.

13. An apparatus comprising: a memory storing a plurality of matrix elements in a linear format; as well as A plurality of computing units, each of the computing units comprising one or more single instruction multiple data (SIMD) units, the plurality of computing units being configured to: generating a request to convert the plurality of matrix elements from the linear format to a tiled format in which the matrix elements of a source matrix are stored as a plurality of tiles, wherein each of the plurality of tiles has fewer elements than the source matrix; When two or more computing units among the plurality of computing units execute the same kernel in parallel, each of the two or more computing units is configured to: loading a given plurality of matrix elements in the plurality of tiles from the cache in parallel through different channels of the cache; as well as performing a matrix operation on the plurality of matrix elements in the tiled format to generate a plurality of results in the tiled format; as well as Classification of the first data set is achieved based on the plurality of results.

14. The apparatus of claim 13, wherein the plurality of computational units are configured to load matrix elements on the plurality of channels in parallel in a single clock cycle.

15. The apparatus of claim 14, wherein each of the two or more computing units comprises a matrix operation unit having a plurality of channels, wherein there is at least one SIMD unit.

16. The apparatus of claim 13, wherein in response to receiving a request to convert a plurality of matrix elements of a first source matrix from the linear format to a tiled format, the apparatus is configured to: reading values ​​from sequential locations of a first buffer in the memory, wherein the first buffer stores matrix elements in the linear format; and When writing the value to a second buffer, stepping through the second buffer with a stride equal to a tile height, wherein the second buffer stores matrix elements in the tiled format.

17. The apparatus of claim 13, wherein matrix elements are transmitted along a plurality of lanes of each of the plurality of computational units, and wherein the matrix operation is a matrix multiplication operation.

18. The apparatus of claim 13, wherein the source matrix corresponds to a first data set and the classification of the first data set is implemented during execution of a machine learning engine application.

Citation Information

Patent Citations

  • Rearranging data between vector and matrix forms in a SIMD matrix processor

    US20020198911A1