Multivariable strided read operations for accessing matrix operands
Through multivariate span reading operation, matrix operands are directly extracted and converted from memory, solving the problem of inefficiency in the prior art and achieving efficient matrix operation processing.
Patent Information
- Application Number
- CN202510105872.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-29
- Filing Date
- 2020-06-24
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art requires complete reading, conversion of formats or sequences from memory when processing matrix operands, which is inefficient and increases processing delay, memory access delay and memory utilization.
The multivariate step-by-step reading operation is adopted to directly extract the matrix operand from the memory, and the memory address of the data elements in the sub-matrix is determined and read through programming parameters, including step-by-step parameters.
Efficient extraction and conversion of matrix operands is realized, processing delays and memory access delays are reduced, and memory utilization is improved.
Smart Images

Figure CN120010919A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with application date of June 24, 2020, application number 202010589581.2, and name “Multi-variable stride read operation for accessing matrix operands”. Technical Field
[0002] The present disclosure relates generally to the field of matrix processing systems and, more particularly, but not exclusively, to multi-variable strided read operations for fetching matrix operands from memory. Background Art
[0003] Training artificial neural networks and / or performing reasoning using neural networks typically requires many computationally intensive operations involving complex matrix arithmetic (e.g., matrix multiplication and convolution of many large multi-dimensional matrix operands). The memory layout of these matrix operands is very important for the overall performance of the neural network. In some cases, for example, matrix operands stored in memory in a specific format may need to be extracted and / or converted to a different format to perform certain operations on the underlying matrix elements. For example, in order to perform certain neural network operations, the dimensions of the matrix operands may need to be rearranged (shuffled) or reordered, or some parts of the matrix operands may need to be extracted, sliced, pruned, and / or reordered. In many computing architectures, this requires the original matrix operands to be read from the memory in full, converted to the appropriate format or order, stored back in the memory as new matrix operands, and then operated. This approach can be extremely inefficient because it increases processing latency, memory access latency, and memory utilization. Summary of the invention
[0004] According to one aspect of the present application, a device is provided, comprising: a memory for storing a matrix for calculation in a neural network, the matrix comprising one or more sub-matrices; and a memory unit for reading data from the memory in the following manner: obtaining one or more programming parameters, the one or more programming parameters being used to read data elements in the matrix from the memory, the one or more programming parameters comprising a stride parameter, the stride parameter indicating the storage size of a memory segment storing a sub-matrix of the matrix, determining a memory address of one or more data elements in the sub-matrix based on the stride parameter, and reading the one or more data elements from the memory segment based on the memory address.
[0005] According to another aspect of the present application, a method is provided, comprising: storing a matrix for calculation in a neural network in a memory, the matrix comprising one or more sub-matrices; obtaining one or more programming parameters, the one or more programming parameters being used to read data elements in the matrix from the memory, the one or more programming parameters comprising a stride parameter, the stride parameter indicating the storage size of a memory segment storing a sub-matrix of the matrix; determining a memory address of one or more data elements in the sub-matrix based on the stride parameter; and reading the one or more data elements from the memory segment based on the memory address. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The present disclosure is best understood from the following detailed description when read in conjunction with the accompanying drawings. It is emphasized that, in accordance with standard practice in the industry, various features are not necessarily drawn to scale and are for illustration purposes only. Where scale is shown explicitly or implicitly, only one illustrative example is provided. In other embodiments, the dimensions of various features may be arbitrarily increased or decreased for clarity of discussion.
[0007] Figure 1 An example embodiment of a matrix processing system that fetches matrix operands from memory utilizing multi-variable strided read operations is illustrated.
[0008] Figure 2 An example of the memory layout of a matrix operand registered with dimensions 71x71, element type BFLOAT16 is shown.
[0009] Figure 3 An example of the memory layout of a matrix operand registered with dimension 33×33 and element type SP3 is shown.
[0010] Figure 4 , 5 6A-6B illustrate examples of multi-variable strided read operations.
[0011] Figure 7 A flow chart illustrating an example embodiment of a multi-variable strided read operation.
[0012] Figures 8A-8B is a block diagram illustrating a general vector friendly instruction format and instruction templates thereof according to an embodiment of the present disclosure.
[0013] Figures 9A-9D is a block diagram illustrating an exemplary specific vector friendly instruction format according to an embodiment of the present disclosure.
[0014] Fig.10 is a block diagram of a register architecture according to one embodiment of the present disclosure.
[0015] Fig.11Ais a block diagram illustrating both an exemplary in-order pipeline and an exemplary register renaming, out-of-order issue / execution pipeline according to embodiments of the present disclosure.
[0016] Fig. 11B is a block diagram illustrating both an exemplary embodiment of an in-order architecture core and an exemplary register renaming, out-of-order issue / execution architecture core to be included in a processor according to embodiments of the present disclosure.
[0017] Figures 12A-12B A block diagram illustrating a more specific exemplary sequential core architecture, where a core is one of multiple logic blocks in a chip (including other cores of the same type and / or different types).
[0018] Fig.13 is a block diagram of a processor that may have more than one core, may have an integrated memory controller, and may have integrated graphics according to an embodiment of the present disclosure.
[0019] Fig.14 , 15 , 16 and 17 are block diagrams of exemplary computer architectures.
[0020] Fig.18 is a block diagram of converting binary instructions in a source instruction set into binary instructions in a target instruction set using a software instruction converter according to an embodiment of the present disclosure. Specific embodiments
[0021] The following disclosure provides many different embodiments or examples for realizing the different features of the present disclosure. The following describes specific examples of components and arrangements to simplify the present disclosure. Of course, these are merely examples and are not intended to be limiting. In addition, the present disclosure may repeat reference numerals and / or letters in various examples. This repetition is for the purpose of simplicity and clarity, and does not itself indicate the relationship between the various embodiments and / or configurations discussed. Different embodiments may have different advantages, and any embodiment does not necessarily require a specific advantage.
[0022] Multivariable strided read operations for extracting matrix operands
[0023] Figure 1 An example embodiment of a matrix processing system 100 is illustrated that utilizes a multi-variable stride read operation to extract matrix operands from a memory. In some cases, for example, matrix operands stored in a memory in a particular format may need to be extracted and / or converted to a different format to perform a particular matrix operation. Therefore, in the illustrated embodiment, as further described below, a multi-variable stride read operation may be performed to extract matrix operands directly from a memory in an appropriate format or order required for a particular matrix operation.
[0024] In the illustrated embodiment, the matrix processing system 100 includes a matrix processor 110 communicatively coupled to a host computing device 120 via a communication interface or bus 130. In some embodiments, for example, the matrix processor 110 may be a hardware accelerator and / or a neural network processor (NNP) designed to accelerate artificial intelligence (AI), machine learning (ML), and / or deep learning (DL) functions on behalf of the host computing device 120, which may be a general-purpose computing platform (e.g., an Intel Xeon platform) that runs artificial intelligence applications. For example, in some cases, the matrix processor 110 may be used to train an artificial neural network and / or perform inference on behalf of the host computing device 120, which typically requires many computationally intensive operations to be performed using complex matrix arithmetic (e.g., matrix multiplications and convolutions of many large multi-dimensional matrix operands).
[0025] The memory layout of these matrix operands or tensors is very important to the overall performance of the neural network. In some cases, for example, matrix operands stored in memory in a particular format may need to be extracted and / or converted to a different format to perform certain operations on the underlying matrix elements. For example, in order to perform certain neural network operations, the dimensions of the matrix operands may need to be rearranged or reordered, or certain parts of the matrix operands may need to be extracted, sliced, pruned, and / or reordered.
[0026] In the illustrated embodiment, for example, multi-dimensional matrix operands 102 may be transferred from a host computing device 120 to a matrix processor 110 in a particular format using direct memory access (DMA) (e.g., via a PCIe interface 130), but the matrix processor 110 may subsequently need to transform the matrix operands 102 into another format 104 to perform certain operations.
[0027] As an example, for a convolutional layer of a neural network, matrix operands used to represent image data may be organized in the base memory layout as dimensions of order CHW x N, where:
[0028] C = channel depth (e.g., the number of channels in each input image);
[0029] H = height (e.g., the height of each input image);
[0030] W = width (eg, the width of each input image); and
[0031] N = number of input images (eg, mini-batch size).
[0032] However, based on the layer description and the operations to be performed, the matrix processor 110 may often need to rearrange or reorder the dimensions of the matrix operands before operating on them. For example, if the matrix operands reside in memory in CHW x N format, the matrix processor 110 may need to perform dimension rearrangement to convert the operands to NHW x C format for a particular layer of the neural network.
[0033] As another example, matrix operands used to represent a set of convolutional filters may be organized in the base memory layout as dimensions of order CRS x K, where:
[0034] C = channel depth;
[0035] R = filter height;
[0036] S = filter width; and
[0037] K = number of filters.
[0038] In some cases, the matrix processor 110 may require certain portions of the convolution filters stored in memory in the CRS x K format to be extracted, sliced, trimmed, and / or reordered into another format.
[0039] In many computing architectures, performing operations on modified matrix operands requires reading the original matrix operands in their entirety from memory, converting them to the appropriate format or order, storing them back into memory as new matrix operands, and then performing the operations. This process can be very inefficient because it increases processing latency, memory access latency, and memory utilization.
[0040] Thus, in the illustrated embodiment, the matrix processing system 100 is capable of operating on matrix operands of various formats in a more flexible and efficient manner. Specifically, the matrix processor 110 can perform a multi-variable stride (MVS) read operation to directly fetch the matrix operands 104 from the memory in the appropriate format or order required for a particular matrix operation. In this manner, matrix operands stored in memory out of order 102 can be fetched and / or converted to the correct format or order 104 using the MVS read operation.
[0041] In some embodiments, for example, an MVS operation can read matrix operands from memory using a software programmable stride read sequence. Specifically, the stride read sequence can be programmed using a sequence of stride read operations (e.g., stride and band operations) that are designed to read matrix operands from memory in an appropriate format or order.
[0042] For example, the stride operation can be used to read memory at a specified stride or offset relative to a base memory address, such as + n strides from the previously read memory address:
[0043] stride + n = read memory [previous address + n]
[0044] The banding operation may be used to read memory at a specified number of sequential memory addresses relative to a base memory address, such as n sequential memory addresses after a previously read memory address:
[0045] Band+n = Read Memory [Previous Address+1]
[0046] Read memory [previous address + 2]
[0047] Read memory [previous address + ...]
[0048] Read memory [previous address + n]
[0049] In this manner, a stride read sequence can be programmed with the appropriate sequence of stride and band operations to read the matrix operands from memory in a desired format or order. This software programmable approach provides great flexibility to the application to extract matrix operands from memory in a variety of different formats in an efficient manner.
[0050] In the illustrated embodiment, for example, the matrix processor 110 includes a controller 112, a matrix element storage and slicing (MES) engine 114, and an execution engine 118. The controller 112 receives instructions from a host computing device 120 and causes the instructions to be executed by the matrix processor 110.
[0051] The matrix element storage and slicing (MES) engine 114 handles the storage and retrieval of matrix operands used by the instructions. Specifically, the MES engine 114 includes an operand memory 115 and a sequence memory 116. The operand memory 115 is used to store matrix data and operands, and the sequence memory 116 is used to store the stride read sequence used to extract the matrix operands from the operand memory 115. In some embodiments, for example, the operand memory 115 may include on-package DRAM (e.g., high bandwidth memory (HBM)) and / or on-chip SRAM (e.g., memory resource blocks (MRBs) of a particular computing cluster), etc.
[0052] The execution engine 118 handles the execution of the instructions. In some embodiments, for example, the execution engine 118 may include one or more matrix processing units (MPUs) designed to perform matrix arithmetic on matrix operands, such as matrix multiplication, convolution, element-by-element arithmetic and logic (e.g., addition (+), subtraction (-), multiplication (*), division ( / ), bitwise logic (AND, OR, XOR, left / right shift), comparison (>, <, >=, <=, ==, !=)), and column-level, row-level, and matrix-wide operations (e.g., sum, maximum, minimum), etc.
[0053] In some embodiments, in order for the matrix processor 110 to execute an instruction that operates on a multi-variable (MV) matrix operand (e.g., a matrix operand to be accessed using an MVS read operation), the following steps are performed as a prerequisite: (i) the MV operand is registered to the matrix processor 110; and (ii) the stride read sequence required to access the MV operand is programmed into the sequence memory 116. In some embodiments, for example, the matrix processor 110 may support certain instructions (e.g., instructions that the host computing device 120 may issue to the matrix processor 110) for performing these steps, such as a register multi-variate operand instruction (REGOP_MV) and a sequence memory write instruction (SEQMEM_WRITE).
[0054] For example, a register multi-variable operand instruction (REGOP_MV) may specify various parameters that collectively identify and / or define the MV operand, such as a handle identifier (ID), a starting memory address (e.g., a base memory address + offset in the operand memory 115), dimensions, a numeric element type (e.g., BFLOAT16 or SP32), and various stride access parameters (e.g., the address of the stride read sequence in the sequence memory 116, the read operation size, the loop size, the super stride), etc. Based on these parameters, the matrix processor 110 assigns the specified handle ID to the MV operand, which enables subsequent instructions to operate on the MV operand by referencing its assigned handle ID.
[0055] Additionally, a sequence memory write instruction (SEQMEM_WRITE) may be used to program a strided read sequence of MV operands into the sequence memory 116 .
[0056] Subsequent instructions that operate on the MV operand can then be executed by referencing the appropriate handle ID. In some embodiments, for example, upon receiving an instruction that references the handle ID assigned to the MV operand, the MV operand is retrieved from operand memory 115 based on the previously registered parameters specified in the REGOP_MV instruction and the corresponding stride read sequence programmed into sequence memory 116, and then the appropriate operation is performed on the MV operand.
[0057] Additional functions and embodiments related to multi-variable stride read operations are further described in conjunction with the remaining figures. Therefore, it should be understood that any aspect of the functions and embodiments described throughout the present disclosure can be implemented Figure 1 A matrix processing system 100 is provided.
[0058] Registration and storage of matrix operands
[0059] In some embodiments, Figure 1 The matrix processor 110 is designed to execute instructions that operate on two-dimensional (2D) matrix operands, each of which is identified by a corresponding "handle". The handle serves as a pointer to a specific 2D matrix operand in memory. In addition, matrix operands with more than two dimensions (e.g., three-dimensional (3D) and two-dimensional (4D) operands) can be supported by organizing the original dimensions into two subsets of dimensions (each subset is considered a single dimension). For example, a 4D matrix operand of dimensions C x H x W x N (channels, height, width, number of images) can be organized as a 2D matrix operand of CHW x N, where dimensions C, H, and W are collectively considered the first dimension and dimension N is considered the second dimension.
[0060] Furthermore, in some embodiments, the matrix processor 110 may support multiple numeric formats for representing the base elements of matrix operands, such as a 16-bit BFLOAT16 or BF16 or a 32-bit single precision floating point format (SP32).
[0061] In some embodiments, handles to matrix operands are registered by software before any instructions to operate on the matrix operands are issued to the matrix processor 110. For example, software executing on the host computing device 120 may register handles to the matrix operands by issuing a register operand instruction to the matrix processor 110.
[0062] Furthermore, in some embodiments, the matrix processor 110 may support multiple variations of register operand instructions, depending on how particular matrix operands are stored in memory. For example, a REGOP instruction may be used to register ordinary matrix operands that are stored in order or sequentially in memory, while a REGOP_MV instruction may be used to register multi-variable (MV) strided matrix operands that are stored out of order in memory.
[0063] For example, for ordinary matrix operands that are stored sequentially or sequentially in memory, the REGOP instruction may include various parameters for registering the matrix operands, such as a handle ID, dimensions (e.g., the size of the x and y dimensions), a starting memory address, and a numeric element type (e.g., BFLOAT16 or SP32):
[0064] REGOP handleID(sizex,sizey)addr ntype
[0065] The registered handle ID is then used as a pointer to the matrix operand in memory. For example, based on the parameters specified in the REGOP instruction, the handle identifier handleID is registered to the hardware to point to a matrix stored in memory at a specific starting memory address (addr), the matrix has dimensions sizex by sizey, and contains a total of sizex*sizey elements of ntype type (e.g., BFLOAT16 or SP32).
[0066] Figure 2-3 An example of a memory layout for various matrix operands registered using a REGOP instruction according to certain embodiments is illustrated. In some embodiments, for example, a memory for storing matrix operands may have a specific depth and width, such as a depth of 4,000 and a width of 512 bits (or 64 bytes). In this way, the memory may store up to 4,000 rows of matrix elements, each row containing up to 32 elements in BFLOAT16 format (512b per row / 16b per element = 32 elements per row) or 16 elements in SP32 format (512b per row / 32b per element = 16 elements per row). In addition, in some embodiments, matrix operands may be stored in memory blocks of a specific size, such as 2 kilobyte (kB) blocks, each block including 32 rows of memory (e.g., 32 rows * 512 bits per row = 16,384 bits = 2048 bytes ≈ 2kB).
[0067] Figure 2 An example 200 of a memory layout of a registered matrix operand with dimensions 71x71 and element type BFLOAT16 is illustrated. In some embodiments, for example, the matrix operands in example 200 may be registered using the REGOP instruction described above.
[0068] In the illustrated example, a logical view of the matrix operand 210 and a corresponding memory layout 220 are shown. Specifically, the matrix operand 210 is logically arranged as 71 rows and 71 columns (71x71) of BFLOAT16 type elements. In addition, for the purpose of storing the matrix operand 210 in the memory 220, the matrix operand 210 is divided into logical blocks (AI), each logical block containing up to 32 rows and 32 columns of elements (32×32). Since the matrix operand 210 does not contain enough elements to completely fill the logical blocks at the right and lower edges, these blocks are smaller than the maximum logical block size of 32x32. For example, although the sizes of logical blocks A, B, D, and E are all 32x32, logical blocks C, F, G, H, and I are smaller because they fall on the right and lower edges of the matrix operand 210.
[0069] In addition, each logic block (AI) of the matrix operand 210 is stored in a fixed-size block of the memory 220 in sequence. In the illustrated embodiment, for example, each memory block is 2kB in size and includes 32 rows of the memory 220 with a width of 512 bits (or 64 bytes), which means that each row has the ability to store up to 32 elements of type BFLOAT16. Therefore, the size of the physical memory block is equal to the maximum size of the logic block (for example, 32 rows of elements, 32 elements per row). In addition, each logic block (AI) is stored in a separate physical memory block, regardless of whether the logic block completely fills the entire memory block. For example, logic blocks A, B, D, and E each fill the entire 32x32 memory block, while logic blocks C, F, G, H, and I are smaller in size, so each of them only partially fills the 32x32 memory block.
[0070] Figure 3 An example 300 illustrates a memory layout of a matrix operand registered with dimensions 33x33 and element type SP32. In some embodiments, for example, the matrix operands in example 300 may be registered using the REGOP instruction described above.
[0071] In the illustrated example, a logical view of a matrix operand 310 and a corresponding memory layout 320 are shown. Specifically, the matrix operand 310 is logically arranged as 33 rows and 33 columns (33×33) of SP32 type elements. In addition, for the purpose of storing the matrix operand 310 in the memory 320, the matrix operand 310 is divided into logical blocks (AD), each logical block containing up to 32 rows and 32 columns of elements (32×32). Since the matrix operand 310 does not contain enough elements to completely fill the logical blocks at the right and lower edges, these blocks are smaller than the maximum logical block size of 32x32. For example, although the maximum size of logical block A is 32x32, logical blocks B, C, and D are smaller because they fall on the right and lower edges of the matrix operand 310.
[0072] In addition, the individual logical blocks (AD) of the matrix operand 310 are sequentially stored in fixed-size blocks of the memory 320. In the illustrated embodiment, for example, each storage block has a size of 2 kB and includes 32 rows of the memory 320 with a width of 512 bits (or 64 bytes), which means that each row has the ability to store up to 16 elements of the SP32 type (e.g., 32×16). Therefore, the size of the physical memory block is half the maximum size of the logical block, because the size of the physical memory block is 32x16, while the maximum size of the logical block is 32x32. As a result, some logical blocks will be larger than a single physical memory block, which means that they will have to be stored separately on two physical memory blocks. Therefore, each logical block (AD) is stored in two physical memory blocks, regardless of whether the logical block completely fills these memory blocks.
[0073] For example, the maximum size of logical block A is 32x32, which means it will completely fill two physical memory blocks. Therefore, logical block A is divided into left logical block A of size 32×16 L and right logic block A R , which are stored in separate physical memory blocks.
[0074] The size of logical block B is 32x1, which means it can fit in a single physical memory block. However, logical block B is still stored in memory 320 using two memory blocks, which means that the first memory block is partially filled with logical block B, while the second memory block is empty.
[0075] The size of logical block C is 1x32, which means it is too large to fit in a single physical memory block, because the width of logical block C is 32 elements, while the width of physical memory block is 16 elements. Therefore, logical block C is split into left logical block C of size 1x16 L and right logic block C R, they are each stored in a separate physical memory block that is only partially filled.
[0076] The size of logical block D is 1x1, which means it can fit in a single physical memory block. Nevertheless, logical block D is stored in memory 320 using two memory blocks, which means that the first memory block is partially filled with logical block D, while the second memory block is empty.
[0077] Multivariable strided matrix operands
[0078] As described above, in some cases, matrix operands stored in a memory in a specific format or order may need to be extracted and / or converted to a different format or order to perform a specific matrix operation. Therefore, in some embodiments, the matrix processor 110 may support a "register multi-variable operand" instruction (REGOP_MV) to register a handle ID for a multi-variable (MV) operand stored in memory in a disorderly and / or non-continuous manner. In this way, subsequent instructions can operate on MV operands by simply referencing their assigned handle IDs. For example, upon receiving a subsequent instruction that references the handle ID assigned to the MV operand, the matrix processor 110 automatically extracts the MV operand from the memory in an appropriate format or order using a multi-variable stride (MVS) read operation, and then performs the appropriate operation corresponding to the received instruction on the MV operand.
[0079] In some embodiments, for example, a REGOP_MV instruction for registering an MV operand may include the following fields and / or parameters:
[0080] REGOP_MV handleID(sizex,sizey,rd_offset,operation_size,seqmem_offset,superstride,loop_size)base_addr ntype
[0081] Once the REGOP_MV instruction is executed, the registered handle ID is used as a pointer to the MV matrix operand in memory and further indicates that the operand should be read from memory in a strided manner (e.g., based on the specified parameters in the REGOP_MV instruction and a corresponding software programmable strided read sequence).
[0082] For example, based on the parameters specified in the REGOP_MV instruction, a handle identifier "handleID" is registered to the hardware and points to a matrix stored in memory starting at a specific starting memory address (base_addr+rd_offset), having a size of "sizex" times "sizey", and containing elements of type "ntype" (e.g., BFLOAT16 or SP32). The registered "handleID" also indicates that the matrix is a multi-variable (MV) operand that should be accessed using an MVS read operation because the matrix may be stored in memory out of order and / or non-contiguously.
[0083] A particular MVS read operation for accessing an MV operand involves a specified number (operation_size) of base read operations performed by looping through a strided read sequence of a specified size (loop_size), performing strided and / or striped read operations within the strided read sequence, and applying a "superstride" between each iteration of the strided read sequence (e.g., until the specified number (operation_size) of base read operations have been performed). In addition, the particular strided read sequence used in the MVS read operation is retrieved from a sequence memory at a specified offset (seqmem_offset) that is pre-programmed into the sequence memory by software (e.g., using a SEQMEM_WRITE instruction).
[0084] In this way, the REGOP_MV instruction enables software to inform the hardware that certain matrix operands should be fetched from memory in a strided manner via a particular strided read sequence specified by the software (e.g., thereby allowing software to select which rows and / or columns to read out from a logical tensor stored in memory, in what order, etc.).
[0085] In some embodiments, for example, the sequence memory is used to store a stride read sequence, which can be used to extract MV operands from the memory. For example, a particular stride read sequence in the sequence memory may include a sequence of stride and / or band instructions. The stride instruction is used to perform a single read operation at a particular stride offset n, while the band instruction is used to perform n sequential read operations.
[0086] In some embodiments, the sequence memory has a depth of 256 entries (e.g., memory addresses) and a width of 9 bits (256 entries x 9 bits). In addition, each entry or address of the sequence memory can store a single instruction of a particular stride read sequence, such as a stride instruction or a band instruction. For example, the most significant bit (MSB) of an entry in the sequence memory can indicate whether the particular instruction is a stride instruction or a band instruction, while the remaining bits can indicate the value of the particular stride or band instruction (e.g., an offset n for a stride operation, or a number n of sequential reads for a band operation).
[0087] A more detailed description of the fields of the REGOP_MV instruction is provided below in Table 1 and in the following sections.
[0088] Table 1: Fields of the Register Multiple Variable Operand (REGOP_MV) instruction
[0089]
[0090]
[0091]
[0092] In some embodiments, for example, when an instruction is received that references a registered "handleID" of a multi-variable (MV) matrix operand (e.g., an operand registered via a REGOP_MV instruction), the MV operand is automatically extracted from memory in an appropriate format or order using a multi-variable stride (MVS) read operation (e.g., based on fields specified in the REGOP_MV instruction and a corresponding stride read sequence in sequence memory).
[0093] For example, an MVS read operation for a particular MV operand performs a first read of memory at the memory address indicated by the "base_addr" + "rd_offset" fields of the REGOP_MV instruction. What happens next depends on the contents of the sequence memory. Specifically, the strided read sequence for a particular MV operand is programmed into the sequence memory at the offset specified in the "seqmem_offset" field of the REGOP_MV instruction, and the number of instructions in the strided read sequence is specified by the "loop_size" field.
[0094] Therefore, the hardware will first read the sequence memory at the address given by "seqmem_offset". If MSB[b8] = 1, this is a band instruction and it indicates the number of additional reads that should be performed on sequential memory locations. If MSB[b8] = 0, this is a stride instruction and it indicates the offset that should be added to the current memory address to find the next expected memory location to read.
[0095] The hardware will continue the process by reading the next sequential memory location and applying the same rules as before. If the previous sequential memory instruction was a stride instruction, the next sequential memory instruction is applied immediately. However, if the previous sequential memory instruction was a band instruction, the next sequential memory instruction is applied only after the specified number of band read operations are completed. Therefore, if a stride instruction is followed by a band instruction, the first read operation is performed at the memory address indicated by the stride offset, and then the specified number of band read operations are performed at the sequential memory addresses that immediately follow. For example, a stride +n instruction followed by a band +3 instruction will read from 4 consecutive memory locations: one due to the stride jump to the memory address at offset +n, and then the other three due to the band reads at the memory addresses that immediately follow. However, any sequence of strides and / or band instructions is allowed. For example, there can be one or more consecutive strides instructions and / or one or more consecutive band instructions (although multiple back-to-back band instructions are inefficient because they can easily be combined into a single band instruction).
[0096] After each memory read operation performed for an MV operand (e.g., a read of a memory address containing an element of an MV operand), a counter associated with the "operation_size" field of the REGOP_MV instruction is incremented by 1. Similarly, after each read of an instruction in sequence memory (e.g., a stride or band instruction), a counter associated with the "loop_size" field of the REGOP_MV instruction is incremented by 1. This process continues until one of two events is detected: "loop_size" or "operation_size" is reached.
[0097] If "loop_size" is reached, a superstride event occurs. This means that the superstride is added to the starting memory address (e.g., base_addr+rd_offset+superstride), and the read operation is then performed at the new starting memory address. After the read is complete, the engine will repeat the sequence memory loop again. It starts by reading the first sequence memory instruction again (e.g., at "seqmem_offset" located in the sequence memory) and applying it as before. It continues to read the same sequence memory instruction until "loop_size" is reached a second time or "operation_size" is reached for the first time. If "loop_size" is reached before "operation_size", the process is repeated at the new starting memory address calculated by applying another superstride (e.g., base_addr+rd_offset+2*superstride). When "operation_size" is reached, the MVS read operation is completed and the MV operand is read from memory.
[0098] In addition, in some embodiments, the MVS read operation can have multiple variants or modes: direct and rotated. For both variants, the method of calculating the memory address used to read the MV operand from the memory is the same. The difference lies in whether a rotated memory selection is applied to the rows read from the memory (e.g., based on the "rd_width" field). For example, in direct mode, the rows read from the memory are treated as rows in the resulting MV operand. However, in rotated mode, the rows read from the memory are rotated or transposed into columns in the resulting MV operand.
[0099] As described above, the contents of the sequence memory inform the hardware of the corresponding stride read sequence for accessing a particular MV operand. The sequence memory can be programmed by software via the sequence memory write (SEQMEM_WRITE) instruction. A description of the fields of the SEQMEM_WRITE instruction is provided in Table 2 below.
[0100] Table 2: Fields of the "Sequence Memory Write" (SEQMEM_WRITE) instruction
[0101]
[0102]
[0103] In some cases, this concept of multivariate matrix operands can be used to perform dimension rearrangement operations. For example, the following primitive is usually needed to support all possible dimension rearrangement operations on multidimensional tensors:
[0104] Original function 1: AB x C->BAx C
[0105] Original function 2: AB x C -> AC x B
[0106] Original function 3: AB x C->C x AB
[0107] However, in the described embodiments, dimension rearrangement operations (e.g., AB x C->BC x A) may be performed directly using MVS read operations. For example, multivariable matrix operands may be registered via the REGOP_MV instruction using various parameters and stride read sequences designed to rearrange or reorder the dimensions of the matrix operands stored in memory.
[0108] Figure 4 An example of an MVS read operation 400 is illustrated. Specifically, the illustrated example shows the layout of matrix operands stored in memory 410, and the instructions of a strided read sequence programmed into sequence memory 420 (which, in this example, contains only stride instructions).
[0109] In addition, the illustrated example describes how the MVS read operation 400 reads an MV operand from the memory 410 based on (i) various fields specified in the REGOP_MV instruction and (ii) a corresponding strided read sequence in the sequence memory 420. Specifically, the illustrated example depicts a row of the memory 410 read by the MVS read operation 400 that contains elements of the MV operand.
[0110] For example, the following pseudo code illustrates how an MVS read operation 400 reads an MV operand from memory 410:
[0111]
[0112]
[0113] Figure 5 Another example of an MVS read operation 500 is illustrated. In the example shown, the stride read sequence programmed into the sequence memory 520 includes stride instructions and band instructions.
[0114] The following pseudo code illustrates how an MVS read operation 500 reads an MV operand from memory 510:
[0115]
[0116]
[0117] Figure 6A-6BAnother example of an MVS read operation 600 is illustrated. In the example shown, the MVS read operation 600 is used to extract a portion of a 3D convolution filter 605 stored in a memory 610 in a CRS×K format. Therefore, the MV operand is registered via the REGOP_MV instruction using the following parameters:
[0118] i. Basic memory address = as shown;
[0119] ii. memory offset = 7;
[0120] iii. Operation size = 36 (eg, total number of read operations = number of pixels per row * number of rows (3*12=36));
[0121] iv. Sequence memory offset = as shown;
[0122] v. Loop size = 7 (number of sequence memory instructions, which determines the number of reads per filter); and
[0123] vi. Superstride = R*S.
[0124] Based on these parameters and the strided access sequence programmed into sequence memory 620, MVS read operation 600 will perform three iterations of the read operation to read the MV operand from memory 610:
[0125] 36 reads in total (operation_size) / 12 reads per iteration (1 initial read + 11 reads for the strided access sequence) = 3 iterations.
[0126] The number of iterations (3) corresponds to the number of channels (C).
[0127] Figure 7 Flowchart 700 illustrates an example embodiment of a multivariable stride (MVS) read operation. In some embodiments, for example, Figure 1 The matrix processing system 100 implements the flowchart 700 .
[0128] The flowchart begins at block 702, where an instruction (e.g., a REGOP_MV instruction) is received for registering a handle of a multi-variable strided matrix operand stored in a memory. For example, the REGOP_MV instruction may include various fields that can be used to extract the matrix operand from the memory, such as a handle identifier (ID), a base memory address, a memory offset, a dimension, a numeric element type (e.g., BFLOAT16 or SP32), an operation size, a sequence memory offset, a loop size, and / or a superstride, etc. Based on these fields, a specified handle ID is assigned to the matrix operand, which enables subsequent instructions to operate on the matrix operand by referencing its assigned handle ID.
[0129] Then, the flow chart proceeds to block 704, where a strided read sequence for accessing the matrix operand is programmed into the sequence memory. In some embodiments, for example, one or more sequence memory write instructions (SEQMEM_WRITE) may be received to program the strided read sequence into the sequence memory. The strided read sequence may include a sequence of strided read operations and / or band read operations, which may be programmed into the sequence memory at an offset specified by a "sequence memory offset" field of a REGOP_MV instruction. In addition, the total number of instructions in the strided read sequence may correspond to a "loop size" field of a REGOP_MV instruction.
[0130] The flow chart then proceeds to block 706 where an instruction is received to perform an operation on a strided matrix operand, such as a matrix multiplication, convolution, or memory copy operation. For example, the instruction may reference a handle ID of a matrix operand. Thus, in order to perform the specific operation required by the instruction, the matrix operand is retrieved from memory in an appropriate format or order using an MVS read operation (e.g., based on the fields specified in the REGOP_MV instruction and the corresponding strided read sequence in the sequence memory).
[0131] For example, an MVS read operation reads a specific number of memory rows (e.g., based on an “operation size” field) containing elements of a matrix operand being read by looping through a strided read sequence of a specific size (based on the “loop size” field), performing strided and / or striped read operations in the strided read sequence, and applying “superstrides” between each iteration of the strided read sequence (e.g., until the number of memory rows required by the “operation size” field have been read from memory).
[0132] Therefore, the flow chart proceeds to block 708 to retrieve the strided read sequence of the matrix operand from the sequence memory. For example, a specific number of entries are read from the sequence memory at an offset indicated by the "sequence memory offset" field of the REGOP_MV instruction. The number of entries read from the sequence memory corresponds to the "loop size" field of the REGOP_MV instruction, and each entry may contain instructions for performing a strided read operation or a band read operation.
[0133] The flow chart then proceeds to block 710 where an iteration of the read operation is performed. For example, in a first iteration, a first row of memory is read at a starting memory address corresponding to the "base memory address" + "memory offset" field of the REGOP_MV instruction. A strided read sequence is then performed (e.g., retrieved from a sequence memory at block 708) to read one or more additional rows of memory relative to the starting memory address of the current iteration. For example, one or more strided read operations and / or stripe read operations in a strided read sequence may be performed.
[0134] The flow chart then proceeds to block 712 to determine whether the MVS read operation is complete. For example, if the total number of lines of memory that have been read is equal to the value in the "operation size" field of the REGOP_MV instruction, then the MVS read operation is complete - otherwise, one or more additional iterations of the MVS read operation are required to read the remaining lines of memory.
[0135] For example, if it is determined at block 712 that the MVS read operation has not completed, the flowchart proceeds to block 714 where a superstride is applied to the starting memory address of the previous iteration to calculate a new starting memory address for the next iteration.
[0136] The flow chart then returns to block 710 to perform another iteration of the read operation relative to the new starting memory address. For example, a row of memory is read at the new starting memory address, and then one or more additional rows of memory are read by repeating the strided read sequence relative to the new starting memory address.
[0137] The process is repeated in this manner until the total number of lines of memory that have been read equals the value in the "Operation Size" field of the REGOP_MV instruction.
[0138] When it is determined at block 712 that the MVS read operation is complete, the flowchart proceeds to block 716 where appropriate operations (eg, operations required by the instruction received at block 706) are performed on the matrix operands.
[0139] At this point, the flow chart may be complete. However, in some embodiments, the flow chart may restart and / or certain frames may be repeated. For example, in some embodiments, the flow chart may restart at block 702 to continue registering handles to access stride matrix operands.
[0140] Example Compute Architecture
[0141] The accompanying drawings described throughout the following sections illustrate example implementations of computing systems, architectures, and / or environments that may be used in accordance with the embodiments disclosed herein. In addition, in some embodiments, certain hardware components and / or instructions described throughout this disclosure may be emulated or implemented as software modules (e.g., in the manner described below).
[0142] Example instruction format
[0143] An instruction set may include one or more instruction formats. A given instruction format may define various fields (e.g., the number of bits, the position of bits) to specify the operation to be performed (e.g., opcode) and the operands and / or other data fields (e.g., masks) on which the operation will be performed. Some instruction formats are further subdivided by the definition of instruction templates (or subformats). For example, an instruction template of a given instruction format may be defined as having different subsets of fields of the instruction format (the included fields are generally in the same order, but at least some have different bit positions because fewer fields are included) and / or defined as having given fields with different interpretations. Therefore, each instruction of the ISA is represented using a given instruction format (and, if defined, in a given instruction template of the instruction format) and includes fields for specifying operations and operands. For example, an exemplary ADD instruction has a specific opcode and instruction format, which includes an opcode field that specifies the opcode and an operand field (source 1 / destination and source 2) that selects an operand; and the ADD instruction that appears in the instruction stream will have specific content in the operand field that selects a specific operand. A set of SIMD extensions known as Advanced Vector Extensions (AVX) (AVX1 and AVX2) and using the Vector Extensions (VEX) encoding scheme have been released and / or announced (see, e.g., 64 and IA-32 Architectures Software Developer's Manual, September 2014; see Advanced Vector Extensions Programming Reference, October 2014).
[0144] Embodiments of the instructions described herein may be implemented in different formats. In some embodiments, for example, (one or more) instructions may be embodied in a "universal vector friendly instruction format", as further described below. In other embodiments, this format is not used, but another instruction format is used, however, the following description of write mask registers, various data transformations (swizzle, broadcast, etc.), addressing, etc. is generally applicable to any possible instruction format. In addition, exemplary systems, architectures, and pipelines are described in detail below. Embodiments of (one or more) instructions may be executed on such systems, architectures, and pipelines, but are not limited to those described in detail.
[0145] Generic vector friendly instruction format
[0146] The vector friendly instruction format is an instruction format suitable for vector instructions (eg, there are certain fields specific to vector operations).While embodiments are described that support vector and scalar operations via the vector friendly instruction format, alternative embodiments use only vector operations of the vector friendly instruction format.
[0147] Figures 8A-8B is a block diagram showing a general vector friendly instruction format and instruction templates thereof according to an example of the present invention. Fig. 8A is a block diagram showing a general vector friendly instruction format and a class A instruction template thereof according to an embodiment of the present disclosure; Figure 8B 800 is a block diagram illustrating a generic vector friendly instruction format and its class B instruction templates according to an embodiment of the present disclosure. Specifically, class A and class B instruction templates are defined for the generic vector friendly instruction format 800, both of which include no memory access 805 instruction templates and memory access 820 instruction templates. In the context of the vector friendly instruction format, the term "generic" refers to an instruction format that is not tied to any specific instruction set.
[0148] Although examples of the present disclosure will be described, the vector friendly instruction format supports the following: 64-byte vector operand length (or size) with 32-bit (4-byte) or 64-bit (8-byte) data element width (or size) (thus, a 64-byte vector consists of 16 doubleword-sized elements, or alternatively, consists of 8 quadword-sized elements); 64-byte vector operand length (or size) with 16-bit (2-byte) or 8-bit (1-byte) data element width (or size); 64-byte vector operand length (or size) with 32-bit (4-byte), 64-bit (8-byte), 16-bit (4-byte), 64-bit (8-byte), 16-bit (8-byte) data element width (or size); 28-bit (2-byte) or 8-bit (1-byte) data element width (or size); and 16-byte vector operand lengths (or sizes) with 32-bit (4-byte), 64-bit (8-byte), 16-bit (2-byte) or 8-bit (1-byte) data element width (or size); alternative embodiments may support more, fewer and / or different vector operand sizes (e.g., 256-byte vector operands) with more, fewer or different data element widths (e.g., 128-bit (16-byte) data element width).
[0149] Fig. 8A The Class A instruction templates include: 1) within the no memory access 805 instruction template, the no memory access, full rounding control type operation 810 instruction template and the no memory access, data transformation type operation 815 instruction template are shown; 2) within the memory access 820 instruction template, the memory access, temporary (temporal) 825 instruction template and the memory access, non-temporal 830 instruction template are shown. Figure 8BThe Class B instruction templates include: 1) within the no memory access 805 instruction template, the no memory access, write mask control, partial rounding control type operation 812 instruction template and the no memory access, write mask control, vsize type operation 817 instruction template are shown; 2) within the memory access 820 instruction template, the memory access, write mask control 827 instruction template is shown.
[0150] The generic vector friendly instruction format 800 includes the following: Figures 8A-8B The following fields are listed in the order shown.
[0151] Format field 840 - The specific value in this field (the instruction format identifier value) uniquely identifies the vector friendly instruction format, and therefore identifies that the instruction appears in the vector friendly instruction format in the instruction stream. Therefore, this field is optional and is not required for instruction sets that only have a generic vector friendly instruction format.
[0152] Basic operation field 842 - its content distinguishes different basic operations.
[0153] Register index field 844 - its contents specify the location of the source and destination operands, either in registers or in memory, either directly or through address generation. These include a sufficient number of bits to select N registers from a PxQ (e.g., 32x512, 16x128, 32x1024, 64x1024) register file. While in one embodiment, N may be up to three sources and one destination register, alternative embodiments may support more or fewer source and destination registers (e.g., may support up to two sources, one of which also acts as a destination, may support up to three sources, one of which also acts as a destination, may support up to two sources and one destination).
[0154] Modifier field 846 - its content distinguishes the occurrence of instructions in the generic vector instruction format that specify memory access from the occurrence of instructions that do not specify memory access; that is, distinguishes between the no memory access 805 instruction template and the memory access 820 instruction template. Memory access operations read and / or write to the memory hierarchy (in some cases using values in registers to specify source and / or destination addresses), while non-memory access operations do not do so (e.g., the source and destination are registers). Although in one embodiment, this field also selects between three different ways to perform memory address calculations, alternative embodiments may support more, fewer, or different ways to perform memory address calculations.
[0155] Enhanced operation field 850 - its content distinguishes which of a variety of different operations are to be performed in addition to the basic operation. This field is context specific. In one example of the present disclosure, this field is divided into a class field 868, an alpha field 852, and a beta field 854. The enhanced operation field 850 allows a common group of operations to be performed in a single instruction instead of 2, 3, or 4 instructions.
[0156] Scale field 860 - its content allows the contents of the index field to be scaled for memory address generation (eg, for address generation using 2scale*index+base).
[0157] Displacement field 862A—its contents are used as part of memory address generation (eg, for address generation using 2scale*index+base+displacement).
[0158] Displacement factor field 862B (note that the juxtaposition of displacement field 862A directly on displacement factor field 862B indicates that one or the other is used) - its contents are used as part of the address generation; it specifies the displacement factor that will be scaled by the size of the memory access (N) - where N is the number of bytes in the memory access (e.g., for address generation using 2scale*index+base+scaled displacement). Redundant low-order bits are ignored, so the contents of the displacement factor field are multiplied by the total size of the memory operand (N) to generate the final displacement used to calculate the effective address. The value of N is determined by the processor hardware at runtime based on the full opcode field 874 (described later) and the data manipulation field 854C. The displacement field 862A and the displacement factor field 862B are optional in the sense that, for example, they are not used in the no memory access 805 instruction template and / or different embodiments may implement only one or neither of the two.
[0159] Data element width field 864 - its contents distinguish which of multiple data element widths is to be used (in some embodiments for all instructions; in other embodiments only for some instructions). This field is optional in the sense that, for example, it is not required if only one data element width is supported and / or certain aspects of the opcode are used to support data element width.
[0160] Write mask field 870 - its content controls, on a per data element position basis, whether the data element positions in the destination vector operand reflect the results of the base and enhanced operations. Class A instruction templates support merge write masks, while class B instruction templates support merge and zero write masks. When merged, the vector mask allows any set of elements in the destination to be protected from update during the execution of any operation (specified by the base and enhanced operations); in another embodiment, the old value of each element of the destination where the corresponding mask bit has a 0 value is retained. Conversely, when the zero vector mask allows any set of elements in the destination to be zeroed during the execution of any operation (specified by the base and enhanced operations); in one embodiment, when the corresponding mask bit has a 0 value, the elements of the destination are set to 0. A subset of this functionality is the ability to control the vector length of the operation being performed (i.e., the span from the first to the last element modified); however, the modified elements do not have to be consecutive. Therefore, the write mask field 870 allows partial vector operations, including loads, stores, arithmetic, logical, etc. Although embodiments of the present disclosure are described in which the contents of the write mask field 870 select one of multiple write mask registers that contains the write mask to be used (and thus the contents of the write mask field 870 indirectly identify the mask to be performed), alternative or additional embodiments allow the contents of the write mask field 870 to directly specify the mask to be performed.
[0161] Immediate field 872 - its content allows the specification of an immediate value. This field is optional in the sense that, for example, it is not present in implementations of the generic vector friendly format that do not support immediate values and it is not present in instructions that do not use immediate values.
[0162] Class field 868 - its content distinguishes different classes of instructions. Fig. 8A -B, the contents of this field select between class A and class B instructions. Fig. 8A -B, rounded squares are used to indicate the presence of a specific value in a field (for example, Fig. 8A -Class A 868A and Class B 868B of the class field 868 in B).
[0163] Class A instruction template
[0164] In the case of the non-memory access 805 instruction templates of class A, the alpha field 852 is parsed as the RS field 852A, whose contents distinguish which of the different enhanced operation types are to be performed (e.g., round 852A.1 and data transform 852A.2 are specified for the no memory access, round type operation 810 and no memory access, data transform type operation 815 instruction templates, respectively), while the beta field 854 distinguishes which specified type of operation is to be performed. In the no memory access 805 instruction templates, the scale field 860, displacement field 862A and displacement factor field 862B are not present.
[0165] No memory access instruction templates – full rounding control type operations
[0166] In the no memory access full round control type operation 810 instruction template, the beta field 854 is parsed as a round control field 854A, the contents of which provide static rounding. Although in the described embodiment of the present disclosure, the round control field 854A includes a suppress all floating point exceptions (SAE) field 856 and a round operation control field 858, alternative embodiments may support encoding these concepts into the same field or having only one or the other of these concepts / fields (e.g., only having a round operation control field 858).
[0167] SAE field 856 - its content distinguishes whether exception event reporting is disabled; when the content of the SAE field 856 indicates that suppression is enabled, the given instruction will not report any type of floating point exception flags and will not trigger any floating point exception handler.
[0168] Round operation control field 858 - its contents distinguish which of a set of rounding operations is to be performed (e.g., round up, round down, round toward zero, and round toward nearest). Thus, the round operation control field 858 allows the rounding mode to be changed on a per-instruction basis. In one embodiment where the processor of the present disclosure includes a control register for specifying the rounding mode, the contents of the round operation control field 850 override the register value.
[0169] No memory access instruction templates – data transformation type operations
[0170] In the no memory access data transform type operation 815 instruction template, the beta field 854 is interpreted as a data transform field 854B, the contents of which distinguish which of a plurality of data transforms is to be performed (eg, no data transform, swipe, broadcast).
[0171] In the case of a memory access 820 instruction template of class A, the alpha field 852 is interpreted as an eviction hint field 852B, whose content distinguishes which eviction hint is to be used (in Fig. 8A, temporary 852B.1 and non-temporary 852B.2 are designated for the memory access, temporary 825 instruction template and the memory access, non-temporary 830 instruction template, respectively), and the beta field 854 is interpreted as a data manipulation field 854C, the contents of which distinguish which of a plurality of data manipulation operations (also referred to as primitives) is to be performed (e.g., no manipulation; broadcast; up-conversion of the source; and down-conversion of the destination). The memory access 820 instruction template includes a scale field 860, and optionally includes a displacement field 862A or a displacement factor field 862B.
[0172] Vector memory instructions utilize conversion support to perform vector loads from memory and vector stores to memory. Like regular vector instructions, vector memory instructions transfer data to / from memory on a data element-by-data element basis, with the actual elements transferred being determined by the contents of the vector mask selected as the write mask.
[0173] Memory access instruction templates – for now
[0174] Temporary data is data that is likely to be reused quickly enough to benefit from caching. However, this is a hint and different processors may implement it in different ways, including ignoring the hint completely.
[0175] Memory access instruction templates – non-temporal
[0176] Non-temporal data is data that is unlikely to be reused quickly enough to benefit from being cached in the first level cache, and should be evicted first. However, this is a hint, and different processors may implement it in different ways, including ignoring the hint entirely.
[0177] Class B instruction template
[0178] In the case of instruction templates of class B, the alpha field 852 is interpreted as a write mask control (Z) field 852C, the content of which distinguishes whether the write mask controlled by the write mask field 870 should be merge or zero.
[0179] In the case of the non-memory access 805 instruction templates of class B, a portion of the β field 854 is parsed as an RL field 857A, the contents of which distinguish which of the different enhanced operation types is to be performed (e.g., round 857A.1 and vector length (VSIZE) 857A.2 are specified for the no memory access, write mask control, partial round control type operation 812 instruction template and the no memory access, write mask control, VSIZE type operation 817 instruction template, respectively), while the remainder of the β field 854 distinguishes which specified type of operation is to be performed. In the no memory access 805 instruction template, the scale field 860, displacement field 862A, and displacement factor field 862B are not present.
[0180] In the no memory access, write mask control, partial rounding control type operation 810 instruction template, the remainder of the beta field 854 is interpreted as the rounding operation field 859A and exception event reporting is disabled (the given instruction does not report any type of floating point exception flag and does not trigger any floating point exception handler).
[0181] Round operation control field 859A - As with round operation control field 858, its contents distinguish which of a set of rounding operations (e.g., round up, round down, round toward zero, and round toward nearest) is to be performed. Thus, round operation control field 859A allows the rounding mode to be changed on a per-instruction basis. In one embodiment where the processor of the present disclosure includes a control register for specifying the rounding mode, the contents of round operation control field 850 override the register value.
[0182] In the no memory access, write mask control, VSIZE type operation 817 instruction template, the remainder of the β field 854 is parsed as a vector length field 859B, the contents of which distinguish which of multiple data vector lengths (e.g., 128, 256, or 512 bytes) to perform on.
[0183] In the case of a memory access 820 instruction template of class B, a portion of the beta field 854 is parsed as a broadcast field 857B, the contents of which distinguish whether a broadcast type data manipulation operation is to be performed, and the remainder of the beta field 854 is parsed as a vector length field 859B. The memory access 820 instruction template includes a scale field 860 and optionally a displacement field 862A or a displacement factor field 862B.
[0184] With respect to the generic vector friendly instruction format 800, a full opcode field 874 is shown that includes the format field 840, the base operation field 842, and the data element width field 864. Although one embodiment is shown where the full opcode field 874 includes all of these fields, in embodiments where all of these fields are not supported, the full opcode field 874 includes less than all of these fields. The full opcode field 874 provides an operation code (opcode).
[0185] The enhanced operation field 850, data element width field 864, and write mask field 870 allow these features to be specified on a per-instruction basis in the generic vector friendly instruction format.
[0186] The combination of the write mask field and the data element width field creates typed instructions because they allow the mask to be applied based on different data element widths.
[0187] The various instruction templates found in class A and class B are beneficial in different situations. In some embodiments of the present disclosure, different processors or different cores within a processor may support only class A, only class B, or both classes. For example, a high-performance general purpose out-of-order core intended for general computing may support only class B, a core intended to be used primarily for graphics and / or scientific (throughput) computing may support only class A, and a core intended for both may support both (of course, a core with some mixture of templates and instructions from both classes but not all templates and instructions from both classes is also within the scope of the present disclosure). In addition, a single processor may include multiple cores, all of which support the same class or different cores support different classes. For example, in a processor with separate graphics and general purpose cores, one of the graphics cores intended to be used primarily for graphics and / or scientific computing may support only class A, while one or more general purpose cores may be a high-performance general purpose core with out-of-order execution and register renaming, intended for general purpose computing that supports only class B. Another processor that does not have a separate graphics core may include one or more general purpose sequential or out-of-order cores that support both class A and class B. Of course, in different embodiments of the present disclosure, features from one class may also be implemented in another class. A program written in a high-level language will be placed (e.g., just-in-time or statically compiled) into a variety of different executable forms, including: 1) a form with only instructions of the class supported by the target processor for execution; or 2) a form with alternative routines written using different combinations of instructions of all classes and with control flow code that selects the routine to be executed based on the instructions supported by the processor currently executing the code.
[0188] Exemplary specific vector friendly instruction format
[0189] Fig. 9 is a block diagram showing an exemplary specific vector friendly instruction format according to an embodiment of the present disclosure. Fig. 9 shows a specific vector friendly instruction format 900, which is specific in the following sense, such as the position, size, parsing and order of the format specified fields, and the values of some of these fields. The specific vector friendly instruction format 900 can be used to extend the x86 instruction set, so some fields are similar or identical to the fields used in the existing x86 instruction set and its extensions (such as AVX). This format is consistent with the prefix encoding field, real opcode byte field, MOD R / M field, SIB field, displacement field and immediate field of the existing x86 instruction set with extensions. The fields in Fig. 8 to which the fields in Fig. 9 are mapped are shown.
[0190] It should be understood that although for illustrative purposes, embodiments of the present disclosure are described with reference to the specific vector friendly instruction format 900 in the context of the general vector friendly instruction format 800, the present disclosure is not limited to the specific vector friendly instruction format 900, unless otherwise stated. For example, the general vector friendly instruction format 800 contemplates various possible sizes for various fields, while the specific vector friendly instruction format 900 is shown as having fields of specific sizes. As a specific example, although the data element width field 864 is shown as a one-bit field in the specific vector friendly instruction format 900, the present disclosure is not limited thereto (i.e., the general vector friendly instruction format 800 contemplates data element width fields 864 of other sizes).
[0191] The general vector friendly instruction format 800 includes the following: Fig.9A The following fields are listed in the order shown.
[0192] EVEX prefix (bytes 0-3) 902 - encoded in four-byte form.
[0193] Format field 840 (EVEX byte 0, bits [7:0]) - The first byte (EVEX byte 0) is the format field 840 and it contains 0x62 (a unique value used to distinguish the vector friendly instruction format in one embodiment of the present disclosure).
[0194] The second through fourth bytes (EVEX bytes 1-3) include a number of bit fields that provide specific capabilities.
[0195] REX field 905 (EVEX byte 1, bits [7-5]) - consists of the EVEX.R bit field (EVEX byte 1, bits [7] - R), the EVEX.X bit field (EVEX byte 1, bits [6] - X), and the EVEX.B bit field (857BEX byte 1, bits [5] - B). The EVEX.R, EVEX.X, and EVEX.B bit fields provide the same functionality as the corresponding VEX bit fields and are encoded using 1's complement form, i.e., ZMM0 is encoded as 1111B and ZMM15 is encoded as 0000B. The other fields of the instruction encode the lower three bits of the register index, as is known in the art (rrr, xxx, and bbb), so Rrrr, Xxxx, and Bbbb can be formed by adding EVEX.R, EVEX.X, and EVEX.B.
[0196] REX' Field 810 - This is the first portion of the REX' field 810 and is the EVEX.R' bit field (EVEX byte 1, bit [4] - R') which is used to encode the upper 16 or lower 16 bits of the extended 32 register set. In one embodiment of the present disclosure, this bit and the other bits shown below are stored in a bit-reversed format to distinguish from the BOUND instruction (in the well-known x86 32-bit mode), whose true opcode byte is 62, but the value 11 in the MOD field is not accepted in the MOD R / M field (described below); alternative embodiments of the present disclosure do not store this and the other indicated bits below in an inverted format. The value 1 is used to encode the lower 16 bits of the register. That is, R'Rrrr is formed by combining EVEX.R', EVEX.R, and other RRRs from other fields.
[0197] Opcode map field 915 (EVEX byte 1, bits [3:0] - mmmm) - its contents encode the implied leading opcode byte (0F, 0F 38, or 0F 3).
[0198] Data element width field 864 (EVEX byte 2, bit [7] - W) - represented by the symbol EVEX.W. EVEX.W is used to define the granularity (size) of the data type (32-bit data element or 64-bit data element).
[0199] EVEX.vvvv 920 (EVEX byte 2, bits [6:3] - vvvv) - The effects of EVEX.vvvv may include the following: 1) EVEX.vvvv encodes the first source register operand, specified in inverted (1's complement) form, and is valid for instructions with 2 or more source operands; 2) EVEX.vvvv encodes the destination register operand, specified in 1's complement form for certain vector shifts; or 3) EVEX.vvvv does not encode any operand, the field is reserved, and should contain 1111b. Therefore, EVEX.vvvv field 920 encodes the 4 low-order bits of the first source register specifier stored in inverted (1's complement) form. Depending on the instruction, additional different EVEX bit fields are used to extend the specifier size to 32 registers.
[0200] EVEX.U 868 Class field (EVEX byte 2, bit [2] - U) - if EVEX.U = 0, it indicates Class A or EVEX.U0; if EVEX.U = 1, it indicates Class B or EVEX.U1.
[0201] Prefix encoding field 925 (EVEX byte 2, bits [1:0]-pp) - provides additional bits for the basic operation field. In addition to providing support for legacy SSE instructions in EVEX prefix format, this also has the benefit of compressing the SIMD prefix (rather than requiring a byte to represent the SIMD prefix, the EVEX prefix only requires 2 bits). In one embodiment, to support legacy SSE instructions that use SIMD prefixes (66H, F2H, F3H) in both the legacy format and the EVEX prefix format, these legacy SIMD prefixes are encoded into the SIMD prefix encoding field; and expanded to the legacy SIMD prefix at runtime before being provided to the PLA of the decoder (so the PLA can execute both the legacy and EVEX formats of these legacy instructions without modification). Although newer instructions can directly use the contents of the EVEX prefix encoding field as an opcode extension, some embodiments expand in a similar manner to maintain consistency, but allow these legacy SIMD prefixes to specify different meanings. Alternative embodiments may redesign the PLA to support 2-bit SIMD prefix encodings, so no expansion is required.
[0202] alpha field 852 (EVEX byte 3, bit [7] - EH; also referred to as EVEX.EH, EVEX.rs, EVEX.RL, EVEX.WriteMaskControl, and EVEX.N; also shown with alpha) - As previously described, this field is context specific.
[0203] Beta field 854 (EVEX byte 3, bits [6:4] - SSS, also known as EVEX.s 2-0 ,EVEX.r 2-0 , EVEX.rr1, EVEX.LL0, EVEX.LLB; also shown with β) - As mentioned before, this field is context specific.
[0204] REX' field 810 - This is the remainder of the REX' field and is the EVEX.V' bit field (EVEX byte 3, bit [3] - V') which can be used to encode the upper 16 or lower 16 bits of the extended 32 register set. This bit is stored in a bit-reversed format. A value of 1 is used to encode the lower 16 bits of the register. That is, V'VVVV is formed by combining EVEX.V', EVEX.vvvv.
[0205] Write mask field 870 (EVEX byte 3, bits [2:0] - kkk) - As previously described, its contents specify the index of a register in the write mask register. In one embodiment of the present disclosure, a particular value EVEX.kkk = 000 has special behavior, implying that no write mask is used for a particular instruction (this can be implemented in a variety of ways, including using hardware that is hardwired to all write masks or bypasses the masking hardware).
[0206] The real opcode field 930 (byte 4) is also called the opcode byte. A portion of the opcode is specified in this field.
[0207] The MOD R / M field 940 (byte 5) includes a MOD field 942, a Reg field 944, and an R / M field 946. As previously described, the content of the MOD field 942 distinguishes between memory access and non-memory access operations. The role of the Reg field 944 can be summarized into two cases: encoding a destination register operand or a source register operand, or being treated as an opcode extension and not used to encode any instruction operand. The role of the R / M field 946 may include the following: encoding an instruction operand that references a memory address, or encoding a destination register operand or a source register operand.
[0208] Scale, Index, Base (SIB) Byte (Byte 6) - As previously described, the contents of the scale fields 860, 952 are used for memory address generation. SIB.xxx 954 and SIB.bbb 956 - The contents of these fields have been previously mentioned in relation to register indexes Xxxx and Bbbb.
[0209] Displacement field 862A (bytes 7-10) - When the MOD field 942 contains 10, bytes 7-10 are the displacement field 862A and it works the same as the traditional 32-bit displacement (disp32) and works at byte granularity.
[0210] Displacement Factor Field 862B (Byte 7) - When the MOD field 942 contains 01, byte 7 is the displacement factor field 862B. The location of this field is the same as the location of the traditional x86 instruction set 8-bit displacement (disp8), which works at byte granularity. Because disp8 is sign-extended, it can only address offsets between -128 and 127 bytes; in terms of a 64-byte cache line, disp8 uses 8 bits and can only be set to 4 very useful values - 128, -64, 0, and 64; because a larger range is often required, disp32 is used; however, disp32 requires 4 bytes. Compared to disp8 and disp32, the displacement factor field 862B is a reinterpretation of disp8; when the displacement factor field 862B is used, the actual displacement is determined by the contents of the displacement factor field multiplied by the size (N) of the memory operand access. This type of displacement is called disp8*N. This reduces the average instruction length (a single byte for the displacement, but with a larger range). This compressed displacement is based on the assumption that the effective displacement is a multiple of the granularity of the memory access, and therefore, there is no need to encode the redundant low-order bits of the address offset. That is, the displacement factor field 862B replaces the traditional x86 instruction set 8-bit displacement. Therefore, the displacement factor field 862B is encoded in the same manner as the x86 instruction set 8-bit displacement (so the ModRM / SIB encoding rules are unchanged), with the only exception that disp8 is overloaded to disp8*N. That is, there is no change in the encoding rules or encoding length, but only in the hardware's parsing of the displacement value (the displacement needs to be scaled with the size of the memory operand to obtain a byte-by-byte address offset). The immediate field 872 operates as described above.
[0211] Full opcode field
[0212] Fig. 9B 874. The vector friendly instruction format 900 of FIG. 874 is a block diagram illustrating the fields of a particular vector friendly instruction format 900 that make up the full opcode field 874 according to one embodiment of the present disclosure. Specifically, the full opcode field 874 includes a format field 840, a basic operation field 842, and a data element width (W) field 864. The basic operation field 842 includes a prefix encoding field 925, an opcode map field 915, and a real opcode field 930.
[0213] Register Index Field
[0214] Fig. 9C844, a block diagram showing the fields of a specific vector friendly instruction format 900 that constitute a register index field 844 according to one embodiment of the present disclosure. Specifically, the register index field 844 includes a REX field 905, a REX' field 910, a MODR / M.reg field 944, a MODR / Mr / m field 946, a VVVV field 920, a xxx field 954, and a bbb field 956.
[0215] Enhanced Action Field
[0216] Fig.9D 850 is a block diagram illustrating the fields of a specific vector friendly instruction format 900 that make up the enhanced operation field 850 according to one embodiment of the present disclosure. When the class (U) field 868 contains 0, it indicates EVEX.U0 (class A 868A); when the class (U) field 868 contains 1, it indicates EVEX.U1 (class B 868B). When U=0 and the MOD field 942 contains 11 (indicating a no memory access operation), the alpha field 852 (EVEX byte 3, bit [7]-EH) is interpreted as the rs field 852A. When the rs field 852A contains 1 (round 852A.1), the beta field 854 (EVEX byte 3, bits [6:4]-SSS) is interpreted as the round control field 854A. The round control field 854A includes a one-bit SAE field 856 and a two-bit round operation field 858. When the rs field 852A contains 0 (data transformation 852A.2), the beta field 854 (EVEX byte 3, bits [6:4]-SSS) is interpreted as a three-bit data transformation field 854B. When U=0 and the MOD field 942 contains 00, 01, or 10 (indicating a memory access operation), the alpha field 852 (EVEX byte 3, bits [7]-EH) is interpreted as an eviction hint (EH) field 852B, and the beta field 854 (EVEX byte 3, bits [6:4]-SSS) is interpreted as a three-bit data manipulation field 854C.
[0217] When U=1, the alpha field 852 (EVEX byte 3, bit [7]-EH) is interpreted as the write mask control (Z) field 852C. When U=1 and the MOD field 942 contains 11 (indicating a no memory access operation), a portion of the beta field 854 (EVEX byte 3, bit [4]-S0) is interpreted as the RL field 857A; when the RL field 857A contains 1 (round 857A.1), the remainder of the beta field 854 (EVEX byte 3, bits [6-5]-S2-1) is interpreted as the round operation field 859A, and when the RL field 857A contains 0 (VSIZE 857.A2), the remainder of the beta field 854 (EVEX byte 3, bits [6-5]-S2-1) is interpreted as the vector length field 859B (EVEX byte 3, bits [6-5]-L1-0). When U=1 and the MOD field 942 contains 00, 01, or 10 (indicating a memory access operation), the beta field 854 (EVEX byte 3, bits [6:4] - SSS) is parsed into a vector length field 859B (EVEX byte 3, bits [6-5] - L1-0) and a broadcast field 857B (EVEX byte 3, bits [4] - B).
[0218] Exemplary Register Architecture
[0219] Fig.10 is a block diagram of a register architecture 1000 according to one embodiment of the present disclosure. In the illustrated embodiment, there are 32 vector registers 1010 that are 512 bits wide; these registers are referenced as zmm0 to zmm31. The low-order 256 bits of the lower 16 zmm registers are overlaid on registers ymm0-16. The low-order 128 bits of the lower 16 zmm registers (the low-order 128 bits of the ymm registers) are overlaid on registers xmm0-15. The specific vector friendly instruction format 900 operates on these overlaid register files as shown in the following table.
[0220]
[0221]
[0222] That is, the vector length field 859B selects between a maximum length and one or more other shorter lengths, where each such shorter length is half the length of the previous length; instruction templates without a vector length field 859B operate on the maximum vector length. In addition, in one embodiment, the class B instruction templates of the specific vector friendly instruction format 900 operate on packed or scalar single / double precision floating point data and packed or scalar integer data. Scalar operations are operations performed on the lowest order data element position in the zmm / ymm / xmm register; according to an embodiment, the high order data element position remains the same as before the instruction or is reset to zero.
[0223] Write mask registers 1015 - In the illustrated embodiment, there are eight write mask registers (k0 through k7), each 64 bits in size. In an alternative embodiment, write mask registers 1015 are 16 bits in size. As previously mentioned, in one embodiment of the present disclosure, vector mask register k0 cannot be used as a write mask; when the encoding that normally represents k0 is used for a write mask, it selects a hardwired write mask of 0xFFFF, effectively disabling write masking for that instruction.
[0224] General Purpose Registers 1025 - In the embodiment shown, there are 16 64-bit general purpose registers that are used with existing x86 addressing modes to address memory operands. These registers are referenced by the names RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 through R15.
[0225] Scalar floating point stack register file (x87 stack) 1045 with MMX packed integer flat register file 1050 aliased on it - in the illustrated embodiment, the x87 stack is an eight-element stack used to perform scalar floating point operations on 32 / 64 / 80-bit floating point data using the x87 instruction set extensions; while the MMX registers are used to perform operations on 64-bit packed integer data, as well as to hold operands for certain operations performed between the MMX and XMM registers.
[0226] Alternative embodiments of the present disclosure may use wider or narrower registers. Additionally, alternative embodiments of the present disclosure may use more, fewer, or different register files and registers.
[0227] Exemplary Core Architectures, Processors, and Computer Architectures
[0228] Processor cores can be implemented in different ways, can be implemented for different purposes, and can be implemented in different processors. For example, implementations of such cores can include: 1) general-purpose sequential cores for general-purpose computing; 2) high-performance general-purpose out-of-order cores for general-purpose computing; 3) special-purpose cores used primarily for graphics and / or scientific (throughput) computing. Different processor implementations can include: 1) CPUs that include one or more general-purpose sequential cores intended for general-purpose computing and / or one or more general-purpose out-of-order cores for general-purpose computing; 2) coprocessors that include one or more special-purpose cores used primarily for graphics and / or science (throughput). Such different processors result in different computer system architectures, which may include: 1) a coprocessor on a different chip from the CPU; 2) a coprocessor on a separate die in the same package as the CPU; 3) a coprocessor on the same die as the CPU (in which case such a coprocessor is sometimes referred to as dedicated logic (e.g., integrated graphics and / or scientific (throughput) logic) or as a dedicated core); 4) a system on a chip, which may include the described CPU (sometimes referred to as application core(s) or application processor(s), the coprocessor described above, and additional functionality on the same die. An exemplary core architecture is described next, followed by an exemplary processor and computer architecture.
[0229] In-order and out-of-order core block diagram
[0230] Fig.11A is a block diagram illustrating both an exemplary in-order pipeline and an exemplary register renaming, out-of-order issue / execution pipeline according to embodiments of the present disclosure. Fig. 11B is a block diagram illustrating an exemplary embodiment of both an in-order architecture core and an exemplary register renaming, out-of-order issue / execution architecture core to be included in a processor in accordance with an embodiment of the present disclosure. Fig.11A The solid line boxes in -B show the in-order pipeline and in-order core, while the optionally added dashed line boxes show the register renaming, out-of-order issue / execution pipeline and core. Assuming that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.
[0231] exist Fig.11A , the processor pipeline 1100 includes a fetch stage 1102, a length decode stage 1104, a decode stage 1106, an allocation stage 1108, a rename stage 1110, a schedule (also known as dispatch or issue) stage 1112, a register read / memory read stage 1114, an execute stage 1116, a write back / memory write stage 1118, an exception handling stage 1122, and a commit stage 1124.
[0232] Fig. 11BA processor core 1190 is shown, which includes a front end unit 1130 coupled to an execution engine unit 1150, and both the execution engine unit 1150 and the front end unit 1130 are coupled to a memory unit 1170. The core 1190 can be a reduced instruction set computing (RISC) core, a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. Alternatively, the core 1190 can be a special-purpose core, for example, a network or communication core, a compression engine, a coprocessor core, a general-purpose computing graphics processing unit (GPGPU) core, a graphics core, etc.
[0233] The front end unit 1130 includes a branch prediction unit 1132 coupled to an instruction cache unit 1134, the instruction cache unit 1134 is coupled to an instruction translation lookaside buffer (TLB) 1136, the instruction translation lookaside buffer 1136 is coupled to an instruction fetch unit 1138, and the instruction fetch unit 1138 is coupled to a decode unit 1140. The decode unit 1140 (or decoder) can decode instructions and generate one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals as outputs, which are decoded from the original instructions or otherwise reflect or are derived from the original instructions. The decode unit 1140 can be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLA), microcode read-only memories (ROMs), etc. In one embodiment, the core 1190 includes a microcode ROM or other medium that stores microcode for certain macroinstructions (e.g., in the decode unit 1140 or within the front end unit 1130). Decode unit 1140 is coupled to rename / allocator unit 1152 in execution engine unit 1150 .
[0234] The execution engine unit 1150 includes a rename / allocator unit 1152, which is coupled to a retirement unit 1154 and a set of one or more scheduler units 1156. The (one or more) scheduler units 1156 represent any number of different schedulers, including reservation stations, central instruction windows, etc. The (one or more) scheduler units 1156 are coupled to (one or more) physical register file units 1158. Each physical register file unit 1158 represents one or more physical register files, different ones of which store one or more different data types, such as scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points, status (e.g., an instruction pointer that is the address of the next instruction to be executed), etc. In one embodiment, the physical register file units 1158 include vector register units, write mask register units, and scalar register units. These register units can provide architectural vector registers, vector mask registers, and general registers. The physical register file(s) unit(s) 1158 overlap with the retirement unit(s) 1154 to illustrate various ways in which register renaming and out-of-order execution may be implemented (e.g., using reorder buffer(s) and retirement register file(s); using future file(s), history buffer(s), and retirement register file(s); using register map and register pool; etc.). The retirement unit(s) 1154 and the physical register file(s) unit(s) 1158 are coupled to execution cluster(s) 1160. The execution cluster(s) 1160 include a set of one or more execution units 1162 and a set of one or more memory access units 1164. The execution units 1162 may perform various operations (e.g., shifts, additions, subtractions, multiplications) on various types of data (e.g., scalar floating point, packed integer, packed floating point, vector integer, vector floating point). While some embodiments may include multiple execution units dedicated to specific functions or sets of functions, other embodiments may include only one execution unit or multiple execution units that all perform all functions.Scheduler unit(s) 1156, physical register file unit(s) 1158, and execution cluster(s) 1160 are shown as possibly multiple because certain embodiments create separate pipelines for certain types of data / operations (e.g., scalar integer pipelines, scalar floating point / packed integer / packed floating point / vector integer / vector floating point pipelines, and / or memory access pipelines, each of which has its own scheduler unit, physical register file unit, and / or execution cluster - and in the case of a separate memory access pipeline, certain embodiments where only the execution cluster of that pipeline has memory access unit(s) 1164 are implemented). It should also be understood that where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution while the rest are in-order issue / execution.
[0235] The set of memory access units 1164 is coupled to a memory unit 1170, which includes a data TLB unit 1172 coupled to a data cache unit 1174, which is coupled to a level 2 (L2) cache unit 1176. In an exemplary embodiment, the memory access unit 1164 may include a load unit, a store address unit, and a store data unit, each of which is coupled to the data TLB unit 1172 in the memory unit 1170. The instruction cache unit 1134 is also coupled to a level 2 (L2) cache unit 1176 in the memory unit 1170. The L2 cache unit 1176 is coupled to one or more other levels of cache and ultimately to the main memory.
[0236] As an example, the exemplary register renaming out-of-order issue / execution core architecture can implement pipeline 1100 as follows: 1) instruction fetch 1138 performs fetch and length decode stages 1102 and 1104; 2) decode unit 1140 performs decode stage 1106; 3) rename / allocator unit 1152 performs allocation stage 1108 and rename stage 1110; 4) (one or more) scheduler units 1156 perform scheduling stage 1112; 5) (one or more) physical register file units 1158 and memory units 1170 perform register read / memory read stage 1114; execution cluster 1160 performs execution stage 1116; 6) memory unit 1170 and (one or more) physical register file units 1158 perform write back / memory write stage 1118; 7) various units may be involved in exception handling stage 1122; 8) retirement unit 1154 and (one or more) physical register file units 1158 perform commit stage 1124.
[0237] Core 1190 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions that have been added with newer versions); the MIPS instruction set from MIP Technologies of Sunnyvale, California; the ARM instruction set from ARM Holdings of Sunnyvale, California (with optional additional extensions, such as NEON)), including the instruction(s) described herein. In one embodiment, core 1190 includes logic to support packed data instruction set extensions (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be executed using packed data.
[0238] It should be understood that a core may support multithreading (executing two or more sets of operations or threads in parallel) and may do so in a variety of ways, including time-sliced multithreading, simultaneous multithreading (where a single physical core provides a logical core for each thread that the physical core is multithreading simultaneously), or a combination thereof (e.g., time-sliced fetch and decode and thereafter simultaneous multithreading, e.g., in Hyper-Threading Technology).
[0239] Although register renaming is described in the context of out-of-order execution, it should be understood that register renaming can be used in a sequential architecture. Although the embodiment of the processor shown also includes separate instruction and data cache units 1134 / 1174 and a shared L2 cache unit 1176, alternative embodiments may have a single internal cache for both instructions and data, such as a level 1 (L1) internal cache, or multiple levels of internal cache. In some embodiments, the system may include a combination of internal caches and external caches, where the external cache is external to the core and / or processor. Alternatively, all caches may be external to the core and / or processor.
[0240] Specific Exemplary In-order Core Architecture
[0241] Fig. 12A -B shows a block diagram of a more specific exemplary sequential core architecture, where a core would be one of several logic blocks in a chip (possibly including other cores of the same type and / or different types). The logic block communicates with some fixed function logic, memory I / O interfaces, and other necessary I / O logic via a high bandwidth interconnect network (e.g., a ring network), depending on the application.
[0242] Fig. 12A1 is a block diagram of a single processor core and its connection to an on-die interconnect network 1202 and its local subset in a level 2 (L2) cache 1204 according to an embodiment of the present disclosure. In one embodiment, the instruction decode unit 1200 supports the x86 instruction set with a packed data instruction set extension. The L1 cache 1206 allows low latency access to cache memory into the scalar and vector units. Although in one embodiment (to simplify the design), the scalar unit 1208 and the vector unit 1210 use separate register sets (scalar registers 1212 and vector registers 1214, respectively), and data transferred between them is written to memory and then read back from the level 1 (L1) cache 1206, alternative embodiments of the present disclosure may use different approaches (e.g., using a single register set or including a communication path that allows data to be transferred between two register files without being written and read back).
[0243] The local subset 1204 of the L2 cache is part of the global L2 cache, which is divided into separate local subsets, one local subset for each processor core. Each processor core has a direct access path to its own local subset 1204 of the L2 cache. The data read by the processor core is stored in its L2 cache subset 1204 and can be quickly accessed in parallel with other processor cores that access their own local L2 cache subsets. The data written by the processor core is stored in its own L2 cache subset 1204 and is flushed from other subsets when necessary. The ring network ensures the consistency of shared data. The ring network is bidirectional to allow agents such as processor cores, L2 caches, and other logic blocks to communicate with each other within the chip. Each ring data path is 1012 bits wide in each direction.
[0244] Fig. 12B According to the embodiment of the present disclosure Fig. 12A An expanded view of a portion of a processor core in FIG. Fig. 12B Includes the L1 data cache 1206A portion of the L1 cache 1206, as well as more details about the vector unit 1210 and vector registers 1214. Specifically, the vector unit 1210 is a 16-wide vector processing unit (VPU) (see 16-wide ALU 1228) that executes one or more of integer, single-precision floating point, and double-precision floating point instructions. The VPU supports swizzling of register inputs through the swizzle unit 1220, digital conversion using digital conversion units 1222A-B, and copying of memory inputs using the copy unit 1224. Write mask registers 1226 allow vector writes to be predicted.
[0245] Fig.13is a block diagram of a processor 1300 that may have more than one core, may have an integrated memory controller, and may have integrated graphics according to an embodiment of the disclosure. Fig.13 The solid-line box in the figure shows a processor 1300 having a single core 1302A (and corresponding cache unit 1304A), a system agent 1310, and a set of one or more bus controller units 1316; but the optional addition of the dashed-line box shows an alternative processor 1300 having the following: multiple cores 1302A-N (and corresponding cache units 1304A-N), a set of one or more integrated memory controller units 1314 in the system agent unit 1310, and dedicated logic 1308.
[0246] Thus, different implementations of processor 1300 may include: 1) a CPU with dedicated logic 1308 (where the dedicated logic is integrated graphics and / or scientific (throughput) logic (which may include one or more cores)), and cores 1302A-N (which are one or more general purpose cores (e.g., general purpose sequential cores, general purpose out-of-order cores, or a combination of both); 2) a coprocessor with cores 1302A-N (where cores 1302A-N are a large number of dedicated cores primarily used for graphics and / or scientific (throughput)); 3) a coprocessor with cores 1302A-N (where cores 1302A-N are a large number of general purpose sequential cores). Thus, processor 1300 may be a general purpose processor, a coprocessor, or a special purpose processor, such as a network or communication processor, a compression engine, a graphics processor, a GPGPU (general purpose graphics processing unit), a high throughput many integrated core (MIC) coprocessor (including 30 or more cores), an embedded processor, and the like. The processor may be implemented on one or more chips. Processor 1300 may be part of and / or may be implemented on one or more substrates using any of a variety of process technologies (eg, BiCMOS, CMOS, or NMOS).
[0247] The memory hierarchy includes one or more levels of cache within the core, a set or one or more shared cache units 1306, and external memory (not shown) coupled to the set of integrated memory controller units 1314. The set of shared cache units 1306 may include one or more mid-level caches (e.g., level 2 (L2), level 3 (L3), level 4 (L4)), or other levels of cache, last level cache (LLC), and / or combinations thereof. While in one embodiment, a ring-based interconnect unit 1312 interconnects the integrated graphics logic 1308, the set of shared cache units 1306, and the system agent unit 1310 / (one or more) integrated memory controller units 1314, alternative embodiments may interconnect these units using any number of well-known techniques. In one embodiment, coherency is maintained between the one or more cache units 1306 and the cores 1302A-N.
[0248] In some embodiments, one or more of the cores 1302A-N may be capable of multithreading. System agent 1310 includes those components that coordinate and operate cores 1302A-N. System agent unit 1310 may include, for example, a power control unit (PCU) and a display unit. The PCU may be or may include logic and components required to regulate the power state of cores 1302A-N and integrated graphics logic 1308. The display unit is used to drive one or more externally connected displays.
[0249] The cores 1302A-N may be homogeneous or heterogeneous in terms of the architectural instruction set; that is, two or more of the cores 1302A-N may be capable of executing the same instruction set, while other cores may be capable of executing only a subset of the instruction set or a different instruction set.
[0250] Exemplary Computer Architecture
[0251] Figure 14-17 is a block diagram of an exemplary computer architecture. Other system designs and configurations known in the art for laptop computers, desktop computers, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and various other electronic devices are also suitable. Generally, a variety of systems or electronic devices that can be combined with the processors and / or other execution logic disclosed herein are generally suitable.
[0252] Reference now Fig.14, a block diagram of a system 1400 according to one embodiment of the present disclosure is shown. The system 1400 may include one or more processors 1410, 1415 coupled to a controller hub 1420. In one embodiment, the controller hub 1420 includes a graphics memory controller hub (GMCH) 1490 and an input / output hub (IOH) 1450 (which may be on separate chips); the GMCH 1490 includes a memory and graphics controller coupled to a memory 1440 and a coprocessor 1445; the IOH 1450 couples an input / output (I / O) device 1460 to the GMCH 1490. Alternatively, one or both of the memory and graphics controller are integrated within the processor (as described herein), the memory 1440 and coprocessor 1445 are directly coupled to the processor 1410, and the controller hub 1420 in a single chip with the IOH 1450.
[0253] Fig.14 The optional nature of the additional processor 1415 is indicated by dashed lines in FIG. 14. Each processor 1410, 1415 may include one or more of the processing cores described herein, and may be some version of processor 1300.
[0254] The memory 1440 may be, for example, dynamic random access memory (DRAM), phase change memory (PCM), or a combination of the two. For at least one embodiment, the controller hub 1420 communicates with the processor(s) 1410, 1415 via a multi-drop bus, such as a front-side bus (FSB), a point-to-point interface such as a QuickPath Interconnect (QPI), or a similar connection 1495.
[0255] In one embodiment, coprocessor 1445 is a special purpose processor, such as a high throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, etc. In one embodiment, controller hub 1420 may include an integrated graphics accelerator.
[0256] There may be various differences between the physical resources 1410 , 1415 in terms of the range of metrics including architectural characteristics, microarchitectural characteristics, thermal characteristics, power consumption characteristics, and the like.
[0257] In one embodiment, processor 1410 executes instructions that control general types of data processing operations. Embedded within the instructions may be coprocessor instructions. Processor 1410 recognizes these coprocessor instructions as being of a type that should be executed by the attached coprocessor 1445. Therefore, processor 1410 issues these coprocessor instructions (or control signals representing coprocessor instructions) onto a coprocessor bus or other interconnect to coprocessor 1445. Coprocessor(s) 1445 accept and execute the received coprocessor instructions.
[0258] Reference now Fig.15 , a block diagram of a first more specific exemplary system 1500 according to an embodiment of the present disclosure is shown. Fig.15 As shown, multiprocessor system 1500 is a point-to-point interconnect system and includes a first processor 1570 and a second processor 1580 coupled via a point-to-point interconnect 1550. Each of processors 1570 and 1580 may be a version of processor 1300. In one embodiment of the present disclosure, processors 1570 and 1580 are processors 1410 and 1415, respectively, and coprocessor 1538 is coprocessor 1445. In another embodiment, processors 1570 and 1580 are processor 1410 and coprocessor 1445, respectively.
[0259] Processors 1570 and 1580 are shown as including integrated memory controller (IMC) units 1572 and 1582, respectively. Processor 1570 also includes point-to-point (PP) interfaces 1576 and 1578 as part of its bus controller unit; similarly, second processor 1580 includes PP interfaces 1586 and 1588. Processors 1570, 1580 can exchange information via point-to-point (PP) interface 1550 using PP interface circuits 1578, 1588. Fig.15 As shown, IMCs 1572 and 1582 couple the processors to respective memories (ie, memory 1532 and memory 1534), which may be portions of main memory locally attached to the respective processors.
[0260] Processors 1570, 1580 may each exchange information with chipset 1590 via respective PP interfaces 1552, 1554 using point-to-point interface circuits 1576, 1594, 1586, 1598. Chipset 1590 may optionally exchange information with coprocessor 1538 via high-performance interfaces 1592, 1539. In one embodiment, coprocessor 1538 is a special-purpose processor, such as a high-throughput MIC processor, a network or communication processor, a compression and / or decompression engine, a graphics processor, a GPGPU, an embedded processor, etc.
[0261] A shared cache (not shown) may be included in either processor, or external to both processors but connected to the processors via the PP interconnect, such that in the event that a processor enters a low power mode, local cache information of either or both processors may be stored in the shared cache.
[0262] Chipset 1590 may be coupled to first bus 1516 via interface 1596. In one embodiment, first bus 1516 may be a peripheral component interconnect (PCI) bus, or a bus such as a PCI Express bus or another third generation I / O interconnect bus, but the scope of the present disclosure is not limited in this regard.
[0263] like Fig.15 As shown, various I / O devices 1514 may be coupled to the first bus 1516, and a bus bridge 1518 coupling the first bus 1516 to a second bus 1520. In one embodiment, one or more additional processors 1515 (e.g., a coprocessor, a high throughput MIC processor, a GPGPU, an accelerator (e.g., a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array, or any other processor) are coupled to the first bus 1516. In one embodiment, the second bus 1520 may be a low pin count (LPC) bus. In one embodiment, various devices may be coupled to the second bus 1520 including, for example, a keyboard and / or mouse 1522, communication devices 1527, and storage 1528 such as a disk drive or other mass storage device (which may include instructions / code and data 1530). Additionally, an audio I / O 1524 may be coupled to the second bus 1520. Note that other architectures are possible. For example, instead of Fig.15 Instead of a point-to-point architecture, the system can implement a multi-drop bus or other such architecture.
[0264] Reference now Fig.16 , shows a block diagram of a second more specific exemplary system 1600 according to an embodiment of the present disclosure. Fig.15 and 16 Like elements in the drawings have like reference numerals, and Fig.15 Some aspects of Fig.16 Omitted to avoid ambiguity Fig.16 other aspects.
[0265] Fig.16 It is shown that processors 1570, 1580 may include integrated memory and I / O control logic ("CL") 1572 and 1582, respectively. Thus, CL 1572, 1582 includes an integrated memory controller unit and includes I / O control logic. Fig.16It is shown that not only the memory 1532, 1534 is coupled to the CL 1572, 1582, but also the I / O device 1614 is coupled to the control logic 1572, 1582. The legacy I / O device 1615 is coupled to the chipset 1590.
[0266] Reference now Fig.17 , shows a block diagram of a SoC 1700 according to an embodiment of the present disclosure. Fig.13 Similar elements in the FIG. have similar reference numerals. In addition, the dashed boxes are optional features on more advanced SoCs. Fig.17 , the interconnect unit(s) 1702 are coupled to the following: an application processor 1710, which includes a set of one or more cores 202A-N and a shared cache unit(s) 1306; a system agent unit 1310; a bus controller unit(s) 1316; an integrated memory controller unit(s) 1314; a set or one or more coprocessors 1720, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random access memory (SRAM) unit 1730; a direct memory access (DMA) unit 1732; and a display unit 1740 for coupling to one or more external displays. In one embodiment, the coprocessor(s) 1720 include a special purpose processor, such as a network or communication processor, a compression engine, a GPGPU, a high throughput MIC processor, an embedded processor, etc.
[0267] The embodiments of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present disclosure may be implemented as a computer program or program code executed on a programmable system, the programmable system comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0268] Program code (e.g. Fig.15 The code 1530 shown in ) can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the present application, the processing system includes any system with a processor, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC) or a microprocessor.
[0269] The program code can be implemented in a high-level procedural or object-oriented programming language to communicate with the processing system. If necessary, the program code can also be implemented in assembly language or machine language. In fact, the mechanism described herein is not limited to the scope of any particular programming language. In any case, the language can be a compiled or parsed language.
[0270] One or more aspects of at least one embodiment may be implemented through representative instructions stored on a machine-readable medium, which represents various logic within the processor and, when read by a machine, causes the machine to construct logic to perform the techniques described herein. Such representations, known as "IP cores," may be stored on a tangible machine-readable medium and provided to various customers or manufacturing facilities to be loaded into manufacturing machines that actually manufacture the logic or processor.
[0271] Such machine-readable storage media may include, but are not limited to, a non-transitory tangible arrangement of an article manufactured or formed by a machine or device, including storage media such as a hard disk, any other type of disk, including floppy disks, optical disks, compact disk read-only memory (CD-ROM), compact disk rewritable (CD-RW) and magneto-optical disks, semiconductor devices such as read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM), phase change memory (PCM), magnetic or optical cards, or any other type of medium suitable for storing electronic instructions.
[0272] Therefore, embodiments of the present disclosure also include non-transitory tangible machine-readable media containing instructions or containing design data, such as hardware description language (HDL), which defines the structures, circuits, devices, processors and / or system features described herein. These embodiments may also be referred to as program products.
[0273] Emulation (including binary conversion, code deformation, etc.)
[0274] In some cases, an instruction converter may be used to convert instructions from a source instruction set to a target instruction set. For example, the instruction converter may convert (e.g., using static binary translation, including dynamic binary translation for dynamic compilation), deform, emulate, or otherwise convert instructions to one or more other instructions to be processed by the core. The instruction converter may be implemented in software, hardware, firmware, or a combination thereof. The instruction converter may be on the processor, off the processor, or partially on the processor and partially off the processor.
[0275] Fig.18 1 is a block diagram of using a software instruction converter to convert binary instructions in a source instruction set into binary instructions in a target instruction set according to an embodiment of the present disclosure. In the illustrated embodiment, the instruction converter is a software instruction converter, but alternatively, the instruction converter can be implemented in software, firmware, hardware, or various combinations thereof. Fig.18It is shown that a program in a high-level language 1802 can be compiled using an x86 compiler 1804 to generate x86 binary code 1806, which can be natively executed by a processor 1816 having at least one x86 instruction set core. A processor 1816 having at least one x86 instruction set core represents any processor that can perform substantially the same functions as an Intel processor having at least one x86 instruction set core, thereby achieving substantially the same results as an Intel processor having at least one x86 instruction set core, by performing the following operations: compatibly executing or otherwise processing (1) a large portion of the instruction set of an Intel x86 instruction set core or (2) an object code version of an application or other software targeted to run on an Intel processor having at least one x86 instruction set core. The x86 compiler 1804 represents a compiler that can be operated to generate x86 binary code 1806 (e.g., object code), wherein the binary code can be executed on a processor 1816 having at least one x86 instruction set core with or without additional link processing. Similarly, Fig.18 It is shown that a program in a high-level language 1802 can be compiled using an alternative instruction set compiler 1808 to generate an alternative instruction set binary code 1810 that can be executed natively by a processor 1814 that does not have at least one x86 instruction set core (e.g., a processor having a core that executes the MIPS instruction set of MIPS Technologies of Sunnyvale, California, USA and / or executes the ARM instruction set of ARM Holdings of Sunnyvale, California, USA). The instruction converter 1812 is used to convert the x86 binary code 1806 into code that can be executed natively by the processor 1814 that does not have an x86 instruction set core. The converted code is unlikely to be the same as the alternative instruction set binary code 1810 because it is difficult to make an instruction converter that can implement it; however, the converted code will complete the general operation and consist of instructions from the alternative instruction set. Therefore, the instruction converter 1812 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device that does not have an x86 instruction set processor or core to execute the x86 binary code 1806 through simulation, emulation, or any other process.
[0276] Numerous other changes, substitutions, variations, alterations, and modifications may be ascertained by those skilled in the art, and the present disclosure encompasses all such changes, substitutions, variations, alterations, and modifications that fall within the scope of the appended claims.
[0277] Example Implementations
[0278] The following examples relate to embodiments described throughout this disclosure.
[0279] One or more embodiments may include a matrix processor comprising: a memory for storing matrix operands and a stride read sequence, wherein: the matrix operands are stored in the memory out of order, and the stride read sequence includes a read operation sequence for reading the matrix operands from the memory in a correct order; a control circuit for receiving a first instruction to be executed by the matrix processor, wherein the first instruction is used to instruct the matrix processor to perform a first operation on the matrix operands; a read circuit for reading the matrix operands from the memory based on the stride read sequence; and an execution circuit for executing the first instruction by performing the first operation on the matrix operands.
[0280] In an example embodiment of the matrix processor, the sequence of read operations includes one or more of: a strided read operation for reading memory at a strided memory address, wherein the strided memory address is offset by a stride offset relative to a previous memory address; or a striped read operation for reading memory at one or more sequential memory addresses following the previous memory address.
[0281] In an example embodiment of a matrix processor, a read circuit for reading a matrix operand from a memory based on a strided read sequence is further configured to: read the matrix operand from the memory via multiple iterations of a read operation, wherein: each iteration of the read operation starts at a corresponding starting memory address of the memory; the strided read sequence is at least partially performed in each iteration of the read operation; and the corresponding starting memory address is incremented by a super stride offset between multiple iterations of the read operation.
[0282] In an example embodiment of the matrix processor, the read circuit for reading the matrix operand from the memory via a plurality of iterations of the read operation is further configured to: continuously perform the plurality of iterations of the read operation until a predetermined number of read operations are performed.
[0283] In an example embodiment of the matrix processor: the control circuit is further used to receive a second instruction to be executed by the matrix processor, wherein: the second instruction is to be received before the first instruction; the second instruction is used to instruct the matrix processor to register an identifier of a matrix operand, wherein the identifier is to be registered based on one or more parameters indicating a memory occupancy of the matrix operand in a memory, and wherein the identifier is used to enable the matrix operand to be identified in subsequent instructions; and the first instruction includes a first parameter indicating the identifier of the matrix operand.
[0284] In an example embodiment of the matrix processor, the control circuit is further for receiving a second instruction to be executed by the matrix processor, wherein the second instruction is to be received before the first instruction, and wherein the second instruction is for instructing the matrix processor to program the strided read sequence into the memory.
[0285] In an example embodiment of a matrix processor: a matrix operand includes a plurality of dimensions arranged in a first order; the plurality of dimensions are arranged in a memory in a second order different from the first order; and a strided read sequence is programmed to perform a dimension rearrangement operation to reorder the plurality of dimensions from the second order to the first order.
[0286] In an example embodiment of the matrix processor: matrix operands are stored at a plurality of non-contiguous memory addresses in a memory; and a strided read sequence is programmed to perform a slice operation to fetch the matrix operands from the plurality of non-contiguous memory addresses in the memory.
[0287] In an example embodiment of a matrix processor, the memory comprises: a first memory for storing matrix operands; and a second memory for storing strided read sequences.
[0288] One or more embodiments may include at least one non-transitory machine-accessible storage medium having instructions stored thereon, wherein when the instructions are executed on a machine, the instructions cause the machine to: receive a first instruction to be executed by a matrix processor, wherein the first instruction is used to instruct the matrix processor to perform a first operation on a matrix operand, wherein the matrix operand is stored out of order in a memory of the matrix processor; access a stride read sequence stored in the memory, wherein the stride read sequence includes a read operation sequence for reading the matrix operand from the memory in a correct order; read the matrix operand from the memory based on the stride read sequence; and cause the first instruction to be executed by the matrix processor, wherein the first instruction is to be executed by performing the first operation on the matrix operand.
[0289] In one example embodiment of a storage medium, the sequence of read operations includes one or more of: a strided read operation for reading memory at a strided memory address, wherein the strided memory address is offset by a stride offset relative to a previous memory address; or a striped read operation for reading memory at one or more sequential memory addresses following the previous memory address.
[0290] In one example embodiment of the storage medium, instructions that cause a machine to read a matrix operand from a memory based on a strided read sequence further cause the machine to: read the matrix operand from the memory via multiple iterations of the read operation, wherein: each iteration of the read operation starts at a corresponding starting memory address of the memory; the strided read sequence is at least partially executed in each iteration of the read operation; and the corresponding starting memory address is incremented by a superstride offset between multiple iterations of the read operation.
[0291] In an example embodiment of the storage medium, the instructions that cause the machine to read the matrix operand from the memory via a plurality of iterations of a read operation further cause the machine to: continuously perform the plurality of iterations of the read operation until a predetermined number of read operations are performed.
[0292] In one example embodiment of the storage medium: the instructions further cause the machine to receive a second instruction to be executed by the matrix processor, wherein: the second instruction is to be received before the first instruction; and the second instruction is used to instruct the matrix processor to register an identifier of a matrix operand, wherein the identifier is to be registered based on one or more parameters indicating a memory occupancy of the matrix operand in a memory, and the identifier is used to enable the matrix operand to be identified in subsequent instructions; and the first instruction includes a first parameter indicating an identifier of the matrix operand.
[0293] In an example embodiment of the storage medium, the instructions further cause the machine to receive a second instruction to be executed by the matrix processor, wherein the second instruction is to be received before the first instruction, and wherein the second instruction is to instruct the matrix processor to program the strided read sequence into the memory.
[0294] In an example embodiment of a storage medium: a matrix operand includes a plurality of dimensions arranged in a first order; the plurality of dimensions are arranged in a memory in a second order different from the first order; and a strided read sequence is programmed to perform a dimension rearrangement operation to reorder the plurality of dimensions from the second order to the first order.
[0295] In an example embodiment of the storage medium: the matrix operands are stored at a plurality of non-contiguous memory addresses in the memory; and the strided read sequence is programmed to perform a slice operation to fetch the matrix operands from the plurality of non-contiguous memory addresses in the memory.
[0296] In an example embodiment of a storage medium, a memory includes: a first memory for storing matrix operands; and a second memory for storing strided read sequences.
[0297] One or more embodiments may include a method comprising: receiving a first instruction to be executed by a matrix processor, wherein the first instruction is used to instruct the matrix processor to perform a first operation on a matrix operand, wherein the matrix operand is stored out of order in a memory of the matrix processor; accessing a stride read sequence stored in the memory, wherein the stride read sequence includes a read operation sequence for reading the matrix operand from the memory in a correct order; reading the matrix operand from the memory based on the stride read sequence; and causing the first instruction to be executed by the matrix processor, wherein the first instruction is to be executed by performing the first operation on the matrix operand.
[0298] In an example embodiment of the method, the sequence of read operations includes one or more of: a strided read operation for reading memory at a strided memory address, wherein the strided memory address is offset by a stride offset relative to a previous memory address; or a striped read operation for reading memory at one or more sequential memory addresses following the previous memory address.
[0299] In an example embodiment of the method, reading a matrix operand from a memory based on a strided read sequence includes: reading the matrix operand from the memory via multiple iterations of a read operation, wherein: each iteration of the read operation starts at a corresponding starting memory address of the memory; the strided read sequence is at least partially performed in each iteration of the read operation; and the corresponding starting memory address is incremented by a superstride offset between multiple iterations of the read operation.
[0300] In an example embodiment of the method, the method further comprises: receiving a second instruction to be executed by a matrix processor, wherein: the second instruction is to be received before the first instruction; the second instruction is used to instruct the matrix processor to register an identifier of a matrix operand, wherein the identifier is to be registered based on one or more parameters indicating a memory occupancy of the matrix operand in a memory, and the identifier will be used to enable the matrix operand to be identified in subsequent instructions; and wherein the first instruction comprises a first parameter indicating an identifier of the matrix operand.
[0301] One or more embodiments may include a system comprising: a host processor; and a matrix processor comprising: a memory for storing matrix operands and a stride read sequence, wherein: the matrix operands are stored in the memory out of order, and the stride read sequence includes a read operation sequence for reading the matrix operands from the memory in a correct order; a control circuit for receiving a first instruction to be executed by the matrix processor, wherein the first instruction is used to instruct the matrix processor to perform a first operation on the matrix operands, and wherein the first instruction is to be received from the host processor; a read circuit for reading the matrix operands from the memory based on the stride read sequence; and an execution circuit for executing the first instruction by performing the first operation on the matrix operands.
[0302] In an example embodiment of the system, the sequence of read operations includes one or more of: a strided read operation for reading memory at a strided memory address, wherein the strided memory address is offset by a stride offset relative to a previous memory address; or a striped read operation for reading memory at one or more sequential memory addresses following the previous memory address.
[0303] In an example embodiment of the system, the read circuit that reads the matrix operand from the memory based on the strided read sequence is further used to: read the matrix operand from the memory via multiple iterations of the read operation, wherein: each iteration of the read operation starts at a corresponding starting memory address of the memory; the strided read sequence is at least partially executed in each iteration of the read operation; and the corresponding starting memory address is incremented by a super stride offset between multiple iterations of the read operation.
Claims
1. A device comprising: A memory for storing a matrix for calculation in a neural network, wherein the matrix includes one or more sub-matrices; as well as A memory unit for reading data from the memory by: obtaining one or more programming parameters for reading data elements in the matrix from the memory, the one or more programming parameters including a stride parameter indicating a storage size of a memory segment storing a sub-matrix of the matrix, determining a memory address of one or more data elements in the sub-matrix based on the stride parameter, and The one or more data elements are read from the memory segment based on the memory address.
2. The device according to claim 1, wherein: The one or more data elements in the sub-matrix are sequentially stored in the memory segment, and the one or more programming parameters further include an offset parameter indicating a memory address offset of a first data element stored in the memory segment.
3. The device according to claim 2, wherein: Determining the memory address includes: The memory address is determined based on the stride parameter and the offset parameter.
4. The device according to claim 1, wherein: Determining the memory address includes: The memory address is determined based on the stride parameter and a base memory address.
5. The device according to claim 1, wherein: The stride parameter corresponds to the stride of data elements between consecutive rows of the matrix.
6. The device according to claim 1, wherein: The stride parameter corresponds to the stride of data elements between consecutive columns of the matrix.
7. The device according to claim 1, wherein: The stride parameter is determined based on the number of data elements along a dimension of the matrix.
8. The device according to claim 1, wherein: Determining the memory address includes: A determination is made as to whether a read cycle including one or more read operations has completed.
9. The device according to claim 8, wherein: Determining the memory address further includes: After determining that the read cycle has completed, the memory address is determined.
10. The device according to claim 8, wherein: Determining the memory address further includes: After determining that the read cycle is not complete, determining the memory address is suspended.
11. A method comprising: storing in a memory a matrix for computation in a neural network, the matrix comprising one or more sub-matrices; Obtaining one or more programming parameters for reading data elements in the matrix from the memory, the one or more programming parameters including a stride parameter indicating a storage size of a memory segment storing a sub-matrix of the matrix; determining a memory address of one or more data elements in the sub-matrix based on the stride parameter; as well as The one or more data elements are read from the memory segment based on the memory address.
12. The method according to claim 11, wherein: The one or more data elements in the sub-matrix are sequentially stored in the memory segment, and the one or more programming parameters further include an offset parameter indicating a memory address offset of a first data element stored in the memory segment.
13. The method according to claim 12, wherein: Determining the memory address includes: The memory address is determined based on the stride parameter and the offset parameter.
14. The method according to claim 11, wherein: Determining the memory address includes: The memory address is determined based on the stride parameter and a base memory address.
15. The method according to claim 11, wherein: The stride parameter corresponds to the stride of data elements between consecutive rows of the matrix.
16. The method according to claim 11, wherein: The stride parameter corresponds to the stride of data elements between consecutive columns of the matrix.
17. The method according to claim 11, wherein: The stride parameter is determined based on the number of data elements along a dimension of the matrix.
18. The method according to claim 11, wherein: Determining the memory address includes: A determination is made as to whether a read cycle including one or more read operations has completed.
19. The method according to claim 18, wherein: Determining the memory address further includes: After determining that the read cycle has completed, the memory address is determined.
20. The method according to claim 18, wherein: Determining the memory address further includes: After determining that the read cycle is not complete, determining the memory address is suspended.
21. One or more non-transitory computer-readable media storing instructions executable to implement operations comprising: storing in a memory a matrix for computation in a neural network, the matrix comprising one or more sub-matrices; Obtaining one or more programming parameters for reading data elements in the matrix from the memory, the one or more programming parameters including a stride parameter indicating a storage size of a memory segment storing a sub-matrix of the matrix; determining a memory address of one or more data elements in the sub-matrix based on the stride parameter; as well as The one or more data elements are read from the memory segment based on the memory address.
22. The one or more non-transitory computer readable media of claim 21, wherein: The one or more data elements in the sub-matrix are sequentially stored in the memory segment, and the one or more programming parameters further include an offset parameter indicating a memory address offset of a first data element stored in the memory segment.
23. The one or more non-transitory computer readable media of claim 22, wherein: Determining the memory address includes: The memory address is determined based on the stride parameter and the offset parameter.
24. The one or more non-transitory computer readable media of claim 21, wherein: Determining the memory address includes: The memory address is determined based on the stride parameter and a base memory address.
25. The one or more non-transitory computer readable media of claim 21, wherein: The stride parameter corresponds to the stride of data elements between consecutive rows or consecutive columns of the matrix.
26. The one or more non-transitory computer readable media of claim 21, wherein: The stride parameter is determined based on the number of data elements along a dimension of the matrix.
27. The one or more non-transitory computer readable media of claim 21, wherein: Determining the memory address includes: determining whether a read cycle including one or more read operations has completed; After determining that the read cycle has completed, the memory address is determined.
28. An apparatus comprising: A computer processor for executing computer program instructions; as well as a non-transitory computer readable memory storing computer program instructions executable by the computer processor to implement operations comprising: storing in a memory a matrix for computation in a neural network, the matrix comprising one or more sub-matrices; Obtaining one or more programming parameters for reading data elements in the matrix from the memory, the one or more programming parameters including a stride parameter indicating a storage size of a memory segment storing a sub-matrix of the matrix; determining a memory address of one or more data elements in the sub-matrix based on the stride parameter; and The one or more data elements are read from the memory segment based on the memory address.
29. The device according to claim 28, wherein The one or more data elements in the sub-matrix are stored in the memory segment in sequence, and the one or more programming parameters also include an offset parameter, which indicates a memory address offset of a first data element stored in the memory segment, wherein determining the memory address includes: determining the memory address based on the stride parameter and the offset parameter.
30. The device according to claim 28, wherein The stride parameter corresponds to the stride of data elements between consecutive rows or consecutive columns of the matrix.
Citation Information
Cited By
Mask generation method and device, computer equipment, readable storage medium and program product
CN121209959A