An instruction set architecture for matrix operations.
The ISA enhances matrix operation performance by using a Configuration Register to reinterpret vector instructions as matrix instructions, addressing inefficiencies in handling multi-dimensional data sets and improving machine learning application performance.
Patent Information
- Application Number
- JP2024569552
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-26
- Filing Date
- 2023-05-25
- Publication Date
- 2025-06-05
AI Technical Summary
Existing instruction set architectures (ISAs) are inefficient when handling multi-dimensional data sets, such as matrices, due to limited vector register resources, leading to reduced computational performance in machine learning operations.
The proposed ISA introduces a new Configuration Register (CR) for matrix operations, allowing vector multiply instructions to be reinterpreted as matrix multiply operations, effectively treating vector register operands as vectors of small matrices rather than single-element vectors.
This solution significantly enhances computational intensity for matrix operations without modifying existing vector instructions, improving performance in machine learning applications while maintaining backward compatibility.
Smart Images

Figure 2025517518000001_ABST
Abstract
Description
[Background technology]
[0001] This specification relates to computer processors and instruction set architectures. An instruction set architecture (ISA) is a model of the behavior of a particular family of processors, independent of the specific hardware implementation or microarchitecture details of any of the processors in the family. An ISA generally defines what kinds of instructions can be executed, what fields the instructions have, the names of configuration and data registers, the types of data, and other characteristics of the family of processors. An ISA provides an abstraction that allows processors with different physical characteristics and capabilities to run the same software. Thus, hardware that implements an ISA can be upgraded to newer or more powerful versions without changing the software.
[0002] Some ISAs define processor support for vector operations, which operate on vectors of any length and relieve the software developer or compiler from having to explicitly represent iterations over the elements of a vector. Instead, processors that implement the ISA automatically iterate over vectors according to the size of the vector, which can be specified at runtime rather than being hard-coded. Processors that implement such vector instructions often utilize dedicated vector processing hardware components with multiple cores that are used to parallelize the vector operations.
[0003] An ISA that defines vector operations may define a set of special vector registers that are used to support vector operations. Vector instructions may then reference the vector registers as operands. An implementation of a vector operation executes the vector instruction without software specifying explicit iterative instructions. To use such vector operations, software may specify various configuration information about the vector and its elements, such as the number of elements in the vector, the size and type of each element in the vector, etc.
[0004] However, although vector operations offer great flexibility for single-dimensional data sets, such arbitrary-length vector operations tend to be inefficient when dealing with multi-dimensional data sets, such as matrices. One problem is that because matrices have two-dimensional indexing, it is highly likely that the processor can run out of vector register resources when attempting to iterate over a two-dimensional matrix of any size. When this occurs, other mitigation measures must be used that reduce computational performance, such as the slow process of writing data out to memory to free up resources in the vector registers.
[0005] This phenomenon is a significant bottleneck in machine learning operations, which usually require very intensive matrix computations. Summary of the Invention
[0006] This specification describes an Instruction Set Architecture (ISA) having instructions that are particularly useful for and improve the performance of matrix operations and related machine learning applications. To do so, the ISA defines a new Configuration Register (CR) for matrix operations and an accompanying instruction set for setting the value of the CR.
[0007] Setting the value of CR to a matrix operation effectively overrides the meaning of the vector multiply instructions, causing these instructions to cause the processor to perform a matrix multiply operation. In doing so, processors that implement the ISA reinterpret vector register operands as vectors of small matrices rather than single-element vectors. For example, instead of the processor operating on a 256-element vector of scalar values, the processor could reinterpret the data as a vector that is 1 / 4 the length of a 2x2 matrix.
[0008] This arrangement provides significantly higher computational intensity without fundamentally modifying existing vector instructions.
[0009] Certain embodiments of the subject matter described herein may be implemented to achieve one or more of the following advantages: The instruction set architecture described herein improves the performance of processors performing matrix operations, making such processors more efficient and faster at executing machine learning applications that rely on such matrix applications. The matrix extensions are also fully backward compatible, allowing older software written specifically for vector operations to continue to run on new processors that implement the matrix extensions. According to an embodiment, a processor is provided that is configured to implement an instruction set architecture having instructions that, when operational, set a configuration register of the processor with one or more values that cause the processor to reinterpret one or more vector instructions as matrix instructions.
[0010] Matrix extension is self-scalable without the need to implement a processor to use a specific matrix size. Furthermore, in a heterogeneous processing environment with cores of different performance and efficiency, it is conceivable that the cores can support different matrix sizes, as long as the OS is careful not to migrate threads from a high-performance core to a low-performance core during matrix processing.
[0011] The processor may be configured to perform vector operations on sequences of matrices and to reinterpret vector instructions as matrix instructions.
[0012] Reinterpreting the vector instruction as a matrix instruction may include reinterpreting the data in the vector register as a sequence of matrices.
[0013] Reinterpreting the data in the vector register as a sequence of matrices may include reinterpreting the data in the vector register as a sequence of 2x2, 4x4, 8x8, or 16x16 matrices.
[0014] The configuration register may have a field that represents the matrix width. The field representing the matrix width may represent the exponent N of a matrix having a width given by 2^N.
[0015] The configuration register may have a field that indicates row and column data order. The configuration register may have a field that indicates a widening mode.
[0016] The configuration register may have a field representing a horizontal accumulation span, and the processor is configured to interpret the value of the horizontal accumulation span as an instruction to use a pre-add instruction during multiply-accumulate operations.
[0017] The instruction set architecture may specify an enable bit in a second, distinct configuration register that specifies whether the processor interprets one or more vector instructions as referencing vector inputs or matrix inputs.
[0018] According to a further embodiment, there is provided a method performed by a processor implementing an instruction set architecture having instructions for setting a configuration register of the processor that controls whether vector instructions are reinterpreted as matrix instructions, the method including executing the instructions for setting the configuration register, receiving one or more vector instructions, and reinterpreting the one or more vector instructions as matrix instructions based on information set in the configuration register.
[0019] Also provided are one or more computer storage media encoded with instructions of an instruction set architecture having instructions for setting a configuration register to control whether a processor implementing the instruction set architecture reinterprets vector instructions as matrix instructions, the instructions executed by a processor implementing the instruction set architecture causing the processor to perform an operation, the instructions including executing instructions to set the configuration register, receiving one or more vector instructions, and, as a result, reinterpreting the one or more vector instructions as matrix instructions.
[0020] The following optional features may be applied to the above method or computer storage medium. Reinterpreting a vector instruction as a matrix instruction may involve performing a vector operation on a sequence of matrices.
[0021] Reinterpreting the vector instruction as a matrix instruction may include reinterpreting the data in the vector register as a sequence of matrices.
[0022] Reinterpreting the data in the vector register as a sequence of matrices may include reinterpreting the data in the vector register as a sequence of 2x2, 4x4, 8x8, or 16x16 matrices.
[0023] Executing the instruction may set a field in a configuration register that represents a matrix width.
[0024] The field representing the matrix width may represent the exponent N of a matrix having a width given by 2^N.
[0025] Executing the instruction may set a field in a configuration register that represents the row and column data order.
[0026] Executing the instruction may set a field in a configuration register that represents the extended mode.
[0027] Executing the instruction may include setting a field in a configuration register representing a horizontal accumulation span, and further interpreting the value of the horizontal accumulation span as an instruction to use an add-ahead instruction during the multiply-accumulate operation.
[0028] The instruction set architecture may specify an enable bit in a second, distinct configuration register that specifies whether the processor interprets one or more vector instructions as referencing vector inputs or matrix inputs.
[0029] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. [Brief description of the drawings]
[0030] [Figure 1] 1 illustrates an exemplary processor for implementing an exemplary instruction set architecture (ISA). [Figure 2A] 1 illustrates an exemplary interpretation of a matrix multiply instruction. [Figure 2B] 2B illustrates an example operation of the matrix multiply instruction of FIG. 2A. [Figure 2C] 2B illustrates an example result of the matrix instruction of FIG. 2A. [Diagram 3]3 is a flowchart illustrating an example process 300 for reinterpreting a vector instruction as a matrix instruction. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0031] 1 illustrates an exemplary processor 102 for implementing an exemplary instruction set architecture (ISA). The processor 102 includes an instruction decode module 110, a standard processing subsystem 130, a configuration subsystem 120, a vector processing subsystem 140, and a matrix multiplier 150. These are exemplary components that may be used to implement the ISA described herein.
[0032] The processor 102 is configured to implement the ISA described herein. The ISA may include multiple instructions. Each instruction may cause the processor to perform one or more operations. The ISA may have one or more matrix instructions that cause the processor 102 to perform a matrix operation. The ISA may include instructions that set one or more values in the configuration registers 125 of the processor 102 that cause one or more vector instructions to be reinterpreted as matrix instructions. Matrix instructions differ from vector instructions in that the operands of a matrix instruction are two-dimensional data sets and the operands of a vector instruction are one-dimensional data sets.
[0033] The instruction decode module 110 comprises logic circuitry that can decode each of the instructions in the ISA and cause subsystems of the processor 102 to perform the operations necessary to implement the instruction.
[0034] The ISA may have one or more vector instructions that cause the processor 102 to perform vector or matrix operations. The ISA also has instructions that set configuration registers to control such vector or matrix operations. The instruction decode module 110 may route the configuration register instructions to the configuration subsystem 120 and may route the vector instructions to the vector processing subsystem 140. The vector processing subsystem 140 may include one or more vector registers 145 and other suitable hardware for implementing vector instructions. Each vector register may hold data for vector processing.
[0035] A vector instruction is an instruction that causes the processor 102 to perform one or more vector operations. For example, a vadd instruction, when executed by the vector processing subsystem 140, may populate a vector register with an element-by-element addition of two other vector registers. In some implementations, the processor may execute vector instructions using parallel processing hardware. For example, the vector processing subsystem 140 may have an array of processing elements that can perform the operations of a vector add instruction in parallel. Thus, a vector instruction may cause the processor 102 to operate on multiple pairs of data specified by the operands of the instruction. The vector registers 145 may store, for example, one-dimensional arrays of integers, logical values, characters, or floating point numbers, just to name a few. Vector instructions may operate on vectors of any length.
[0036] Vector instructions may include instructions to perform vector operations. In some implementations, vector instructions may reference vector registers 145 as operands. To use such vector operations, configuration registers 125 store data that specifies various configuration information about the vector and its elements, such as the number of elements in the vector, the size and type of each element in the vector, etc.
[0037] For example, the ISA may include instructions to set vector registers with data describing one vector of length M, instructions to set vector registers with data describing M vectors of length 1 to M, and instructions to multiply two vectors. The vector processing subsystem may set operands in a vector register to represent one vector and set operands in another vector register to represent 1 to M vectors. The vector processing subsystem 140 may then multiply the two vectors together.
[0038] The ISA may also have instructions that configure the configuration registers 125 of the processor 102 to reinterpret one or more vector instructions as matrix instructions. A matrix instruction is an instruction that causes the processor to perform operations on two-dimensional data sets of any size. The instruction decode module 110 sends the instructions to the configuration subsystem 120. The configuration subsystem 120 includes one or more configuration registers 125. For one or more of the configuration registers 125, the ISA may define a configuration register (CR) for matrix operations and an associated instruction set for setting the value of the CR.
[0039] Setting the value of CR125 to a matrix operation effectively overrides the meaning of the vector multiply instructions, causing these instructions to cause the processor to perform a matrix multiply operation. In doing so, processors that implement the ISA reinterpret vector register operands as vectors of small matrices, rather than vectors of single elements. For example, instead of the processor operating on a vector of scalar values, the processor could reinterpret the data as a vector that is 1 / 4 the length of a 2x2 matrix.
[0040] An example of a configuration register for matrix operations is now described. The exemplary configuration register has the name vtypex, which has the following fields and abbreviations: select matrix width (vsmw), matrix data order (vmdo), extension mode (vnwmode), and horizontal accumulation span (vhspan).
[0041] The selection matrix width field represents the width of the matrix referenced by the vector instruction. In some implementations, the selection matrix width is specified as the exponent of the expression 2^N. In other words, a value of 0 represents a width of 1, a value of 4 represents a width of 16, and so on. For example, if the vector registers 145 of the processor 102 hold 16 values, then a selection matrix width of 0 is interpreted as 16 scalar values, a selection matrix width of 1 is interpreted as a vector register holding four 2x2 matrices, and a selection matrix width of 2 is interpreted as a vector register holding one 4x4 matrix.
[0042] The matrix data order field specifies whether the values in the vector registers are arranged in row-major or column-major order. This feature effectively provides a free transpose when performing matrix multiplication. In some implementations, the matrix data order field can be set to specify z-ordering or Morton ordering, which effectively interleaves the x and y coordinates.
[0043] The extended mode field specifies the bit width of the computation output. In a normal multiplication operation of two 8-bit numbers, the result can be at most two extended 16-bit numbers. However, 16 bits is often insufficient for machine learning applications that rely on accumulation. Thus, when the extended mode field is set, the processor may allocate more bits to the output result than would normally be the case. Thus, the result of a multiplication of two 8-bit numbers can be stored in a quad-extended 32-bit output register. Conversely, the extended mode field can also be used to narrow the output, for example when the result needs to be shifted and truncated.
[0044] The horizontal accumulation span field affects the operation of the matrix multiplication operation. In effect, this field provides a second addition step after the multiplication but before the accumulation. This feature improves on one drawback of the output quad expansion, which is that the outputs must be written to twice as many output registers as there are inputs, which can be complicated to implement in hardware. Instead, after the multiplication, this field specifies the horizontally scaled sums for groups of matrices, e.g., groups of 2, groups of 4, or groups of 8, thereby reducing the number of outputs that need to be written.
[0045] The ISA can also specify an enable bit (veml) that controls whether a vector instruction is executed in vector mode or matrix mode. In some implementations, the enable bit is the value of a second, different configuration register 125 that controls vector operations. Placing the enable bit in that second register allows for full backward compatibility with prior programs that did not contemplate matrix extensions.
[0046] To set values in the matrix configuration registers, the ISA could define a new instruction to do so, for example vsetvxi. The new instruction could have fields that specify the values to write to the matrix configuration registers, and software could change these values at run time if desired.
[0047] Thus, when a vector operation encounters the enable bit set, the processor 102 treats the input operands as representing a group of matrices rather than a vector of scalars.
[0048] If the enable bits indicate that the vector instruction is being executed in matrix mode, instruction decode module 110 sends the instruction to matrix multiplier 150, which includes appropriate hardware for performing matrix operations on vector register operands using data in vector registers 145, e.g., treating the data in the registers as a sequence of matrices and multiplying the matrices. If the enable bits indicate that the vector instruction is being executed in vector mode, instruction decode module 110 sends the instruction to be executed by vector processing subsystem 140 instead.
[0049] The ISA may also have one or more standard (e.g., non-vector and non-matrix) instructions, such as load, store, add, and branch. The instruction decode module 110 may route the standard instructions to the standard processing subsystem 130, which includes suitable hardware for implementing the standard instructions. For example, the standard processing subsystem 130 may execute a load instruction by issuing a command to memory of data located at a particular address specified by the load instruction.
[0050] 2A illustrates an example interpretation of a matrix multiply instruction. The matrix multiply instruction may be implemented in any suitable processor that implements the ISA described herein, such as processor 102 of FIG.
[0051] In this example, the processor has two vector registers with 16 elements each. A first vector register 210 contains elements V0, V1, V2, ..., V15, and a second vector register 220 contains elements V16, V17, V18, ..., V31. For example, the elements may store data representing integers or floating point numbers.
[0052] With the appropriate configuration registers set, the processor can be configured to interpret instructions in matrix mode instead of vector mode. The processor can be configured to interpret vector register operands as vectors of matrices of a specified size. The processor can reinterpret vector register operands as vectors of matrices of a specified size instead of vectors of single scalar elements. In this example, the processor can reinterpret the data as a vector of 2x2 matrices instead of a length 16 vector of a single element.
[0053] The matrix width may be specified by a mathematical expression. In some implementations, the matrix width is specified as the exponent of the expression 2^N. More specifically, a value of 0 represents a width of 1, a value of 4 represents a width of 16, and so on. In this example, the vector registers hold 16 values. Thus, the select matrix width 1 is interpreted as each vector register holding four 2x2 matrices.
[0054] In this example, the first four elements of the first vector register 210 are interpreted as a 2x2 matrix 212. Each position in the matrix can be represented as (r,c), where r ranges from 0 to row -1 in total, and c ranges from 0 to column -1 in total. In this example, r ranges from 0 to 1, and c also ranges from 0 to 1. The processor interprets matrix 212 as having element V1 at location (0,0), element V2 at location (0,1), element V3 at location (1,0), and element V4 at location (1,1).
[0055] The processor may similarly interpret the remaining elements of the first vector register 210 into three further 2×2 matrices 214 (for elements V4-V7), 216 (for elements V8-V11), and 218 (for elements V12-V15). The processor may also similarly interpret the elements of the second vector register 220 into four 2×2 matrices 222 (for elements V16-V19), 224 (for elements V20-V23), 226 (for elements V24-V27), and 228 (for elements V28-V31).
[0056] In this example, the processor receives a matrix instruction that reads "vmul VR3, VR2, VR1," which can be decoded to instruct the processor to interpret the vector registers as a stored matrix with properties defined by the configuration registers, multiply an element of the first vector register 210 (i.e., VR1) with an element of the second vector register 220 (i.e., VR2), and store the result in the third vector register 230 (i.e., VR3).
[0057] Figure 2B illustrates an example operation of the matrix multiply instruction of Figure 2A. The matrix multiply instruction may be implemented in a processor, such as processor 102 of Figure 1.
[0058] In this example, the processor is configured to interpret instructions in matrix mode, so that the processor can interpret vector registers 210 and 220 as vectors of 2 by 2 matrices. The processor can interpret the matrix instruction "vmul VR3, VR2, VR1" as performing a matrix multiplication between the matrices in the first vector registers 212, 214, 216, and 218 and the matrices in the second vector registers 222, 224, 226, and 228.
[0059] The processor may multiply the first matrix 212 in the first vector register 210 by the first matrix 222 in the second vector register 220. Matrix 212 has V0 at position (0,0) and V1 at position (0,1). V2 is at position (1,0) and V4 is at position (1,1). Matrix 222 has V16 at position (0,0) and V17 at position (0,1). V18 is at position (1,0) and V19 is at position (1,1).
[0060] The result of multiplying 2×2 matrix 212 with 2×2 matrix 222 is another 2×2 result matrix 232. After the matrix multiplication is performed, the (0,0) position of result matrix 232 may contain the result of V0×V16+V1×V18. The (0,1) position of result matrix 232 contains the result of V0×V17+V1×V19. The (1,0) position of result matrix 232 contains the result of V2×V16+V3×V18. The (1,1) position of result matrix 232 contains the result of V2×V17+V3×V19.
[0061] The processor may multiply each remaining matrix in the first vector register 210 with the matrix of the same index in the second vector register 220 to generate a result matrix. Specifically, the processor may multiply the second 2×2 matrix in the first vector register 214 with the second 2×2 matrix in the second vector register 224 to generate a resultant 2×2 matrix 234. Similarly, the processor may multiply matrix 216 with matrix 226 to generate a resultant matrix 236 and matrix 218 with matrix 228 to generate a resultant matrix 238.
[0062] Figure 2C illustrates an example result of the matrix instruction of Figure 2 A. The matrix multiply instruction may be implemented in a processor, such as processor 102 of Figure 1.
[0063] The processor may interpret the matrix instruction “vmul VR3, VR2, VR1” as performing a matrix multiplication between the matrix in the first vector register 210 and the matrix in the second vector register 220, and storing the result in the third vector register 230. The third vector register 230 is of the same dimension as the first vector register 210 and the second vector register 220.
[0064] In this example, the third vector register 230 is a vector of 16 elements. The third vector register 230 stores the matrix values resulting from the vector product operations 232, 234, 236, and 238. The first result matrix 232 is the result of the multiplication of the first 2×2 matrix of the first vector register 210 and the first 2×2 matrix of the second vector register 220. The elements of the first result matrix 232 populate the first four elements of the third vector register 230. Specifically, the first element of the third vector register 230 is the (0,0) index of the first result matrix 232, e.g., V0×V16+V1×V18. The second element of the third register is the (0,1) index of the first result matrix 232, and the third and fourth elements are populated by the (1,0) and (1,1) indices, respectively.
[0065] In this pattern, the elements of the second result matrix 234 populate the fifth through eighth elements of the third vector register 230. The next four elements are populated by elements of the third result matrix 236, and the last four elements are populated by elements of the fourth result matrix 238. The four resulting matrices are thus represented as the third vector register 230.
[0066] 3 is a flow chart illustrating an example process 300 for reinterpreting a vector instruction as a matrix instruction. Process 300 may be performed by a processor, such as processor 102 of FIG.
[0067] The processor executes an instruction that sets a configuration register to reinterpret vector instructions as matrix instructions (step 310). Setting the configuration register to matrix operations effectively overrides the meaning of vector product instructions, thereby causing these instructions to cause the processor to perform matrix multiply operations. In doing so, the processor reinterprets vector register operands as vectors of matrices rather than single-element vectors.
[0068] The configuration instruction may be associated with a matrix width. In some implementations, execution of the instruction sets a field in a configuration register that represents the matrix width. The matrix width field may represent a width of a matrix referenced by the vector instruction. In some implementations, the selected matrix width is specified as an exponent of the expression 2^N.
[0069] The configuration instructions may be associated with a matrix data order. In some implementations, execution of the instructions sets a field in a configuration register that represents the matrix data order. The matrix data order field may specify whether the values in the vector register are arranged in row-major or column-major order. In some implementations, the matrix data order field may be set to specify z-ordering or Morton ordering, which effectively interleaves the x and y coordinates.
[0070] A configuration instruction may be associated with an extended mode. In some implementations, execution of the instruction sets a field in a configuration register that represents the extended mode. The extended mode field may specify a bit width of the computation output. Setting the extended mode field may cause the processor to allocate more bits to the output result. Conversely, the extended mode field may also be used to narrow the output, for example when the result needs to be shifted and truncated.
[0071] The configuration instruction may be associated with a horizontal accumulation span. In some implementations, execution of the instruction sets a field in a register that represents a horizontal accumulation span. The horizontal accumulation span field may affect the matrix multiplication and accumulation operations. In effect, the field specifies to perform a second addition step after the multiplication but before the accumulation. In some examples, execution of the instruction causes the processor to interpret the value of the horizontal accumulation span as a directive to use an add-ahead instruction during the multiply-accumulate operation. The value of the horizontal accumulation span may represent the size of each group of matrices to be input to the add-ahead operation. For example, if the value of the horizontal accumulation span is 2, then each pair of matrices is added together into a single matrix used in the accumulation. The horizontal accumulation span effectively reduces the number of outputs that need to be written during the multiply-accumulate operation.
[0072] The configuration instruction can be associated with an enable bit. In some examples, executing the instruction can specify an enable bit in a second configuration register. The enable bit can specify whether the processor interprets the vector instruction as referencing a vector input of a matrix input.
[0073] The processor receives a vector instruction that references two vector registers (step 320). The vector registers can hold vector data for processing. The vector registers can have a specified number of elements. The vector registers can represent, for example, one-dimensional arrays of integers, logical values, characters, or floating point numbers.
[0074] A vector instruction may cause a processor to perform an operation on two vector registers. For example, a vector instruction may cause a processor to multiply an element of a first vector by an element of the same index of a second vector, e.g., a first element of a first vector register by a first element of a second vector register, a second element of a first vector register by a second element of a second vector register, etc. As another example, a vector instruction may cause a processor to add elements of two vector registers together. In some implementations, a vector instruction may reference more than two vector registers. For example, an instruction may indicate that a result of a multiplication (or addition, etc.) of data in the vector registers should be stored in a third vector register.
[0075] The processor reinterprets the vector instruction as a matrix instruction for matrices stored in two vector registers (step 330). The processor reinterprets the vector register as a vector of matrices of the specified size. For example, if the vector register has 16 elements and the specified size is 2×2, the processor reinterprets the vector register as a vector of four 2×2 matrices. The first element of the vector becomes a matrix containing the first four elements of the original vector register. In some examples, the data in the vector register can be reinterpreted as a sequence of 2×2, 4×4, 8×8, or 16×16 matrices.
[0076] The processor may perform vector operations on a sequence of matrices. For example, if a vector instruction multiplies an element of a first vector by an element at the same index of a second vector, the processor may multiply a first matrix in a first reinterpretation vector register by a first matrix in a second reinterpretation vector register.
[0077] For example, suppose a processor receives a vector multiplication instruction that references two input vectors and a third output vector. If a configuration register specifies that the input is a 2-by-2 matrix, the processor will interpret each successive group of four elements in the input vector register as a 2-by-2 matrix rather than as four scalars, and perform a matrix multiplication with the corresponding group of four values stored in the other input vector register. This strategy can result in significant performance improvements by effectively doubling the performance of each execution lane by reusing each data input twice in the two multiplication operations.
[0078] Certain novel aspects of the subject matter herein are set forth in the following claims. Embodiments of the subject matter and functional operations described herein can be implemented in digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware, such as the structures disclosed herein and their structural equivalents, or one or more combinations thereof. Embodiments of the subject matter described herein can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or for execution to control the operation of a data processing apparatus. Alternatively, or in addition, the program instructions can be encoded in an artificially generated propagated signal, such as a mechanically generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to a receiving apparatus suitable for execution by a data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more thereof. However, the computer storage medium is not a propagated signal.
[0079] The term "data processing apparatus" encompasses any kind of apparatus, device, and machine for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. An apparatus may include special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). In addition to hardware, an apparatus may also include code that creates an environment for the execution of the computer program in question, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof.
[0080] A computer program (which may be referred to or described as a program, software, software application, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted, declarative or procedural, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may correspond to a file in a file system, but this is not necessarily the case. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program, or in multiple coordinated files, e.g., files that store one or more modules, subprograms, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0081] The processes and logic flows described herein may be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may also be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0082] A computer suitable for executing a computer program can include, for example, be based on, a general-purpose or dedicated microprocessor or both, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from a read-only memory, a random access memory, or both. The basic elements of a computer are a central processing unit for implementing and executing instructions, and one or more memory devices for storing instructions and data. Typically, a computer also includes, or is operatively coupled to receive data from, or transfer data to, one or more mass storage devices for storing data, such as, for example, magnetic disks, magneto-optical disks, or optical disks. However, such devices are not required for a computer. Furthermore, a computer can be embedded in other devices, such as, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as, for example, a universal serial bus (USB) flash drive, to name a few.
[0083] Computer-readable media suitable for storing computer program instructions and data include, by way of example, all types of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0084] Although the present specification contains many specific implementation details, these should not be construed as limiting the scope or claimable content of any invention, but as descriptions of features that may be inherent to a particular embodiment of a particular invention. Certain features described in the context of separate embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features of the invention described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, even if features may be described above as functioning in a particular combination and originally claimed as such, one or more features from the claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.
[0085] Similarly, although operations are shown in a particular order in the figures, this should not be understood as requiring that such operations be performed in the particular order or sequential order shown, or that all of the operations shown be performed, to achieve desirable results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the program components and systems described may generally be integrated into a single software product or packaged into multiple software products.
[0086] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. By way of example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.
Claims
1. 1. A processor configured to implement an instruction set architecture having instructions that, when operated, set configuration registers of the processor with one or more values that cause the processor to reinterpret one or more vector instructions as matrix instructions.
2. The processor of claim 1 , wherein the processor is configured to perform vector operations on a sequence of matrices and reinterpret the vector instructions as matrix instructions.
3. The processor of claim 1 , wherein reinterpreting the vector instruction as a matrix instruction comprises reinterpreting data in a vector register as a sequence of matrices.
4. 4. The processor of claim 3, wherein reinterpreting the data in the vector register as a sequence of matrices comprises reinterpreting the data in the vector register as a sequence of 2x2, 4x4, 8x8, or 16x16 matrices.
5. The processor of claim 1 , wherein the configuration register has a field representing a matrix width.
6. 6. The processor of claim 5, wherein the field representing the matrix width represents an exponent N of a matrix having a width given by 2^N.
7. The processor according to claim 1 , wherein the configuration register has a field indicating a row and column data order.
8. The processor according to claim 1 , wherein the configuration register has a field indicating an extended mode.
9. the configuration register having a field representing a horizontal accumulation span; A processor according to any preceding claim, wherein the processor is configured to interpret values in the horizontal accumulation span as instructions to use add-ahead instructions during multiply-accumulate operations.
10. 10. The processor of claim 1, wherein the instruction set architecture specifies an enable bit in a second, different configuration register that specifies whether the processor interprets the one or more vector instructions as referencing vector inputs or matrix inputs.
11. 1. A method performed by a processor implementing an instruction set architecture having instructions for setting a configuration register of the processor that controls whether vector instructions are reinterpreted as matrix instructions, the method comprising: executing the instructions to set the configuration register; receiving one or more vector instructions; and reinterpreting the one or more vector instructions as matrix instructions based on information set in the configuration register; A method comprising:
12. The method of claim 11 , wherein reinterpreting the vector instruction as a matrix instruction comprises performing a vector operation on a sequence of matrices.
13. The method of any one of claims 11 to 12, wherein reinterpreting the vector instruction as a matrix instruction comprises reinterpreting data in a vector register as a sequence of matrices.
14. 14. The method of claim 13, wherein reinterpreting the data in the vector register as a sequence of matrices comprises reinterpreting the data in the vector register as a sequence of 2x2, 4x4, 8x8, or 16x16 matrices.
15. The method of any one of claims 11 to 14, wherein executing the instruction sets a field in the configuration register representing a matrix width.
16. 16. The method of claim 15, wherein the field representing the matrix width represents an exponent N of a matrix having a width given by 2^N.
17. The method of any one of claims 11 to 16, wherein executing the instruction sets a field in the configuration register representing a row and column data order.
18. The method of any one of claims 11 to 17, wherein executing the instruction sets a field in the configuration register that represents an extended mode.
19. Executing the instructions sets a field in the configuration register representing a horizontal accumulation span; The processor of any one of claims 11 to 18, further comprising interpreting a value in the horizontal accumulation span as an instruction to use an add-ahead instruction during a multiply-accumulate operation.
20. 20. The method of any one of claims 11 to 19, wherein the instruction set architecture specifies an enable bit in a second, different configuration register that specifies whether the processor interprets the one or more vector instructions as referencing vector inputs or matrix inputs.
21. One or more computer storage media encoded with instructions of an instruction set architecture having instructions for setting a configuration register to control whether a processor implementing the instruction set architecture reinterprets vector instructions as matrix instructions, The instructions executed by the processor implementing the instruction set architecture cause the processor to perform an operation, the instructions comprising: executing said instructions to set said configuration register; receiving one or more vector instructions; and reinterpreting the one or more vector instructions as matrix instructions based on information set in the configuration register; and reinterpreting the one or more vector instructions as matrix instructions; [0026] In one or more computer storage media,
22. 22. The one or more computer storage media of claim 21, wherein reinterpreting the vector instructions as matrix instructions comprises performing a vector operation on a sequence of matrices.
23. 23. The one or more computer storage media of claim 21, wherein reinterpreting the vector instruction as a matrix instruction comprises reinterpreting data in a vector register as a sequence of matrices.
24. 24. The one or more computer storage media of claim 23, wherein reinterpreting the data in the vector register as a sequence of matrices includes reinterpreting the data in the vector register as a sequence of 2x2, 4x4, 8x8, or 16x16 matrices.
25. 25. The one or more computer storage media of any one of claims 21 to 24, wherein executing the instructions sets a field in the configuration register representing a matrix width.
26. 26. The one or more computer storage media of claim 25, wherein the field representing the matrix width represents an exponent N of a matrix having a width given by 2^N.
27. The one or more computer storage media of any one of claims 21 to 26, wherein executing the instructions sets a field in the configuration register representing a row and column data order.
28. The one or more computer storage media of any one of claims 21 to 27, wherein executing the instructions sets a field in the configuration register representing an extended mode.
29. Executing the instructions sets a field in the configuration register representing a horizontal accumulation span; The one or more computer storage media of any one of claims 21 to 28, further comprising interpreting a value of the horizontal accumulation span as an instruction to use an add-ahead instruction during a multiply-accumulate operation.
30. 30. The one or more computer storage media of any one of claims 21-29, wherein the instruction set architecture specifies an enable bit in a second, different configuration register that specifies whether the processor interprets the one or more vector instructions as referencing vector inputs or matrix inputs.
Citation Information
Patent Citations
Image processor, its method and program
JP2009049460A
Vector processor, computation execution method and program
JP2019159440A
Register-based matrix multiplication
JP2020527778A
Random tag setting instructions for tag-protected memory systems
JP2021517690A
Register addressing information for data transfer instruction
WO2022023701A1