Systems and methods for zeroing a pair of bank registers

Through the matrix (chip) operating system and method, efficient matrix operation is performed using the configured chip, which solves the problem of low efficiency in handling large-scale matrices in the prior art, and realizes efficient calculation of large-scale data.

CN113885942BActive Publication Date: 2025-06-17INTEL CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111269207.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-12-29
Filing Date
2018-11-30
Publication Date
2025-06-17
Estimated Expiration
2038-11-30

AI Technical Summary

Technical Problem

In computing tasks, especially machine learning and batch data processing, it is difficult for the prior art to efficiently perform matrix operations when processing large matrices, especially for larger matrices, the operation efficiency is low.

Method used

By introducing a matrix (chip) operating system and methods, the configured chips are used to perform matrix operations, including matrix multiplication, chip addition, chip subtraction and other operations, and the chip parameters are configured through the TILECONFIG instruction to support efficient matrix operations.

Benefits of technology

It realizes efficient operation of large matrices, improves computing efficiency, and significantly improves the performance of matrix operations, especially when processing large-scale data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113885942B_ABST
    Figure CN113885942B_ABST
Patent Text Reader

Abstract

The embodiments detailed herein relate to systems and methods for zeroing a pair of slice registers. In one example, a processor includes: a decoding circuit for decoding a matrix pair zeroing instruction having a field for an opcode and an identifier for identifying a destination matrix, the destination matrix having a PAIR parameter equal to true; and an execution circuit for executing the decoded matrix pair zeroing instruction to zero each element of a left matrix and a right matrix of the identified destination matrix.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This is a divisional application of a patent application for invention titled "System and Method for Zeroing a Pair of Tile Registers", with an application date of November 30, 2018, a priority date of December 29, 2017, and an application number of 201811452469.3. Technical Field

[0002] The field of the present invention generally relates to computer processor architectures, and more particularly to systems and methods for zeroing a pair of tile registers. Background Art

[0003] Matrices are becoming increasingly important in computing tasks such as machine learning and other bulk data processing. Brief Description of the Drawings

[0004] The present invention is illustrated by way of example and not limitation in the accompanying drawings, in which like reference numerals indicate like elements, wherein:

[0005] Figure 1A Illustrates an embodiment of a configured tile;

[0006] Figure 1B Illustrates an embodiment of a configured tile;

[0007] Figure 2 Illustrates several examples of matrix storage;

[0008] Figure 3 Illustrates an embodiment of a system utilizing a matrix (tile) operation accelerator;

[0009] Figure 4 and Figure 5 Shows different embodiments of how to use a matrix operation accelerator to share memory;

[0010] Figure 6 Illustrates an embodiment of a matrix multiply-accumulate ("TMMA") operation using a tile;

[0011] Figure 7 Illustrates an embodiment of a subset of the iterative execution of a chained fused multiply-accumulate instruction;

[0012] Figure 8 Illustrates an embodiment of a subset of the iterative execution of a chained fused multiply-accumulate instruction;

[0013] Figure 9 Illustrates an embodiment of a subset of the iterative execution of a chained fused multiply-accumulate instruction;

[0014] Figure 10 Illustrates an embodiment of a subset of the iterative execution of a chained fused multiply-accumulate instruction;

[0015] Figure 11 Illustrates a SIMD implementation with a power-of-two size according to an embodiment, where the accumulator uses an input size larger than the size of the input to the multiplier;

[0016] Figure 12 Illustrates an embodiment of a system utilizing a matrix operation circuit;

[0017] Figure 13 Illustrates an embodiment of a processor core pipeline that supports matrix operations using tiles;

[0018] Figure 14 Illustrates an embodiment of a processor core pipeline that supports matrix operations using tiles;

[0019] Figure 15 Illustrates examples of matrices expressed in row-major format and column-major format;

[0020] Figure 16 Illustrates an example of the use of a matrix (tile);

[0021] Figure 17 Illustrates an embodiment of a method of using a matrix (tile);

[0022] Figure 18 Illustrates support for a configuration of the use of tiles according to an embodiment;

[0023] Figure 19 Illustrates an embodiment of a description of supported matrices (tiles);

[0024] Figure 20(A)-Figure 20(D) Illustrates an example of a (multiple) register;

[0025] Figure 21 Illustrates an exemplary execution of a TZPAIR instruction;

[0026] Figure 22 Illustrates an embodiment of a method executed by a processor to process a TZPAIR instruction;

[0027] Figure 23 Illustrates a more detailed description of the execution of a TZPAIR instruction;

[0028] Figure 24 Is exemplary pseudocode describing an embodiment of a method executed by a processor to process a TZPAIR instruction;

[0029] Figure 25A-Figure 25B Is a block diagram illustrating a general vector-friendly instruction format and its instruction template according to an embodiment of the present invention;

[0030] Figure 25Ais a block diagram illustrating a general vector-friendly instruction format and its class A instruction templates according to an embodiment of the present invention;

[0031] Figure 25B is a block diagram illustrating a general vector-friendly instruction format and its class B instruction templates according to an embodiment of the present invention;

[0032] Figure 26A is a block diagram illustrating an exemplary specialized vector-friendly instruction format according to an embodiment of the present invention;

[0033] Figure 26B is a block diagram illustrating a field of a specialized vector-friendly instruction format that constitutes a complete opcode field according to an embodiment of the present invention;

[0034] Figure 26C is a block diagram illustrating a field of a specialized vector-friendly instruction format that constitutes a register index field according to an embodiment of the present invention;

[0035] Figure 26D is a block diagram illustrating a field of a specialized vector-friendly instruction format that constitutes an extended operation field according to an embodiment of the present invention;

[0036] Figure 27 is a block diagram of a register architecture according to an embodiment of the present invention;

[0037] Figure 28A is a block diagram illustrating both an exemplary in-order pipeline and an exemplary register-renamed out-of-order issue / execution pipeline according to an embodiment of the present invention;

[0038] Figure 28B is a block diagram illustrating an exemplary embodiment of an in-order architecture core to be included in a processor and an exemplary register-renamed out-of-order issue / execution architecture core according to an embodiment of the present invention;

[0039] Figure 29A-Figure 29B A block diagram illustrating a more specific exemplary in-order core architecture, which would be one of several logic blocks in a chip (including other cores of the same type and / or different types);

[0040] Figure 29A is a block diagram of a single processor core according to an embodiment of the present invention and its connection to an on-die interconnect network and a local subset of its second-level (L2) cache;

[0041] Figure 29B is according to an embodiment of the present invention Figure 29A an expanded view of a portion of the processor core in;

[0042] Figure 30is a block diagram of a processor that may have more than one core, may have an integrated memory controller, and may have integrated graphics according to an embodiment of the present invention;

[0043] Figure 31-Figure 34 is a block diagram of an exemplary computer architecture;

[0044] Figure 31 A block diagram showing a system according to an embodiment of the present invention;

[0045] Figure 32 is a block diagram of a first more specific exemplary system according to an embodiment of the present invention;

[0046] Figure 33 is a block diagram of a second more specific exemplary system according to an embodiment of the present invention;

[0047] Figure 34 is a block diagram of a system on a chip (SoC) according to an embodiment of the present invention; and

[0048] Figure 35 is a block diagram comparing the use of a software instruction converter to convert binary instructions in a source instruction set into binary instructions in a target instruction set according to an embodiment of the present invention. DETAILED DESCRIPTION

[0049] In the following description, numerous specific details are set forth. However, it should be understood that embodiments of the present invention may be practiced without these specific details. In other instances, well-known circuits, structures, and techniques are not shown in detail to avoid obscuring the understanding of this description.

[0050] References in the specification to "one embodiment," "an embodiment," "an example embodiment," etc. indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment may include the particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an embodiment, it is considered to be within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in conjunction with other embodiments, whether or not explicitly described.

[0051] In many mainstream processors, manipulating matrices is a difficult and / or instruction-intensive task. For example, multiple rows of a matrix can be placed into multiple packed data (e.g., SIMD or vector) registers, and then multiple rows of the matrix can be operated on individually. For example, depending on the data size, adding two 8x2 matrices may require loading or gathering into four packed data registers. Then, perform a first addition on the packed data register corresponding to the first row from each matrix, and perform a second addition on the packed data register corresponding to the second row from each matrix. Then, scatter the resulting packed data registers back to memory. While this scenario may be acceptable for small matrices, it is generally unacceptable for larger matrices.

[0052] I. High-Level Discussion

[0053] Mechanisms for supporting matrix operations in computer hardware such as central processing units (CPUs), graphics processing units (GPUs), and accelerators are detailed herein. Matrix operations utilize two-dimensional (2-D) data representing one or more packed regions of memory (such as registers). Throughout this specification, these 2-D data are referred to as tiles. Note that a matrix can be smaller than a tile (using fewer than all tiles), or can utilize multiple tiles (the matrix is larger than the size of any one tile). Throughout this specification, matrix (tile) language is used to indicate operations performed using the tiles that affect the matrix; whether the matrix is larger than any one tile is generally irrelevant.

[0054] Each tile can be acted upon by different operations, such as those detailed herein, including but not limited to: matrix (tile) multiplication, tile addition, tile subtraction, tile diagonal, tile zero, tile transpose, tile electrode, tile broadcast, tile row broadcast, tile column broadcast, tile multiplication, tile multiply and accumulate, tile shift, etc. Additionally, support for operators such as those using scaling and / or bias, or support for non-numerical applications such as OpenCL, "local memory", data compression / decompression, etc. can be used in the future with these operations.

[0055] Multiple portions of storage (such as (non-volatile and volatile) memory, registers, caches, etc.) are arranged as tiles having different horizontal and vertical scales. For example, a tile can have a horizontal scale of 4 (e.g., four rows of a matrix) and a vertical scale of 8 (e.g., 8 columns of a matrix). Typically, the horizontal scale is the relevant element size (e.g., 2 bits, 4 bits, 8 bits, 16 bits, 32 bits, 64 bits, 128 bits, etc.). Multiple data types (single-precision floating point, double-precision floating point, integer, etc.) can be supported.

[0056] A. Example Use of Configured Tiles

[0057] In some embodiments, tile parameters can be configured. For example, a given tile can be configured to provide tile options. Exemplary tile options include, but are not limited to: the number of rows of the tile, the number of columns of the tile, whether the tile is valid, and whether the tile is composed of pairs of tiles of equal size.

[0058] Figure 1A FIG. illustrates an embodiment of a configured tile. As shown, 4 kB of the application memory 102 has 1 kB tiles stored thereon - tile 0 104, tile 1 106, tile 2 108, and tile 3 110. In this example, these 4 tiles are not composed of pairs, and each tile has elements arranged in rows and columns. Tiles t0104 and t1 106 have K rows and N columns of 4-byte elements (e.g., single-precision data), where K = 8 and N = 32. Tiles t2 108 and t3 110 have K rows and N / 2 columns of 8-byte elements (e.g., double-precision data). Since the width of a double-precision operand is twice that of a single-precision operand, this configuration is consistent with a palette for providing tile options, providing at least 4 kB of total storage to at least 4 names. In operation, tiles can be loaded from and stored to memory using load operations and store operations. Depending on the instruction encoding scheme used, the amount of available application memory and the size, number, and configuration of available tiles vary.

[0059] Figure 1B FIG. illustrates an embodiment of a configured tile. As shown, 4 kB of the application memory 122 can have 2 pairs of 1 kB tiles stored thereon, the first pair being tiles t4L 124 and t4R 126, and the second pair being tiles t5L 128 and t5R 130. As shown, the tile pairs are divided into left tiles and right tiles. In other embodiments, the tile pairs are divided into even tiles and odd tiles. In this example, these 4 tiles each have elements arranged in rows and columns. Tiles t4L 124 and t4R 126 have K rows and N columns of 4-byte elements (e.g., single-precision data), where K = 8 and N = 32. Tiles t5L 128 and t5R 130 have K rows and N / 2 columns of 8-byte elements (e.g., double-precision data). Since the width of a double-precision operand is twice that of a single-precision operand, this configuration is consistent with a palette for providing tile options, providing at least 4 kB of total storage to at least 2 names. Figure 1A The four tiles use 4 names, each name naming a 1 kB tile, while Figure 1B the tile pairs in can use 2 names to specify the paired tiles. In some embodiments, tile instructions accept the names of paired tiles as operands. In operation, tiles can be loaded from and stored to memory using load operations and store operations. Depending on the instruction encoding scheme used, the amount of available application memory and the size, number, and configuration of available tiles vary.

[0060] In some embodiments, tile parameters are configurable. For example, a "palette" is used to provide tile options. Exemplary options include, but are not limited to: the number of tile names, the number of bytes in a stored row, the number of rows and columns in a tile, and so on. For example, the maximum "height" (number of rows) of a tile can be defined as:

[0061] Maximum tile rows = Constructed storage / (Number of palette names / Bytes per row)

[0062] Thus, an application can be written such that the fixed use of names will be able to take advantage of different storage sizes across implementations.

[0063] Configuration of the tile is accomplished using the tile configuration ("TILECONFIG") instruction, where specific tile usage is defined within the selected palette. The declaration includes the number of tile names to be used, the requested number of rows and columns for each name (tile), and in some embodiments, the requested data type for each tile. In some embodiments, a consistency check is performed during the execution of the TILECONFIG instruction to determine that it matches the limitations of the palette entry.

[0064] B. Exemplary Tile Storage Types

[0065] Figure 2 Several examples of matrix storage are illustrated. In (A), the tile is stored in memory. As shown, each "row" consists of four packed data elements. To reach the next "row", a stride value is used. Note that the rows can be stored continuously in memory. When the tile storage does not map to the underlying memory array row width, strided memory access allows access to one row and then to the next row.

[0066] Loading from and storing to memory is typically strided access from the application memory to the packed data rows. Exemplary TILELOAD and TILESTORE instructions or other instructions reference the application memory as the TILE (tile) operand in the load operation instruction and are, in some embodiments, restartable to handle (up to) 2 * row page faults, unmasked floating point exceptions, and / or interrupts per instruction.

[0067] In (B), the matrix is stored in a tile consisting of multiple registers, such as packed data registers (single instruction multiple data (SIMD) or vector registers). In this example, the tile is stacked on three physical registers. Typically, consecutive registers are used; however, this need not be the case.

[0068] In (C), the matrix is stored in a tile of non-register storage that can be accessed by a fused multiply-add (FMA) circuit used in in-tile operations. This storage can be inside the FMA or adjacent to the FMA. Additionally, in some embodiments, as discussed below, this storage can be used for data elements rather than entire rows of the tile.

[0069] Report the supported parameters of the TMMA architecture via CPUID. In some embodiments, the list of information includes the maximum height and the maximum SIMD scale. Configuring the TMMA architecture requires specifying the scale of each tile, the element size of each tile, and the palette identifier. This configuration is done by executing the TILECONFIG instruction.

[0070] Successful execution of the TILECONFIG instruction enables subsequent TILE operators. The TILERELEASEALL instruction clears the tile configuration and disables TILE operations (until the next TILECONFIG instruction is executed). In some embodiments, XSAVE, XSTORE, etc. are used in context switches that use tiles. In some embodiments, 2 XCE0 bits are used in XSAVE, one for TILECONFIF metadata and one bit corresponding to the actual tile payload data.

[0071] TILECONFIG not only configures tile usage but also sets a status variable that indicates that the program is in a code region where the tile is configured. The implementation can enumerate restrictions on other instructions that can be used with the tile region, such as no use of the existing register set, etc.

[0072] Exiting the tile region is typically done using the TILERELEASEALL instruction. This instruction takes no parameters and quickly invalidates all tiles (indicating that the data elements no longer require any save and restore) and clears the internal state corresponding to being in the tile region.

[0073] In some embodiments, the tile operation will zero out any rows and any columns that exceed the scale specified by the tile configuration. For example, as each row is written, the tile operation will zero out the data beyond the configured number of columns (taking into account the element size). For example, for a 64-byte row and a tile configured with 10 rows and 12 columns, the operation of writing FP32 elements will write the output / result data for each of the 10 rows forward by 12 * 4 bytes and zero out the remaining 4 * 4 bytes in each row. The tile operation also completely zeros out any rows after the first 10 configured rows. When using a 1K tile with 64-byte rows, there will be 16 rows, so in this example, the last 6 rows will also be zeroed out.

[0074] In some embodiments, when loading data, context restoration (e.g., XRESTOR) forces data types for lines beyond the configured lines of the tile to be maintained as zero. If there is no valid configuration, all lines are zeroed. XRESTOR for tile data loads junk information in columns beyond those configured columns. It should not be possible to clear XRESTOR beyond the configured number of columns because there is no element width associated with the tile configuration.

[0075] When writing the entire TILE store to memory, context save (e.g., XSAVE) exposes the entire TILE store. If XRESTOR loads junk data into the rightmost part of the tile, that data will be saved by XSAVE. For lines beyond the number specified for each tile, XSAVE will write zero.

[0076] In some embodiments, tile instructions are restartable. Operations accessing memory allow restart after a page fault. Computational instructions handling floating-point operations also allow unmasked floating-point exceptions by masking exceptions controlled by control and / or status registers.

[0077] To support restarting instructions after these events, the instructions store information in the start register detailed below.

[0078] II. Matrix (Tile) Operating System

[0079] A. Exemplary Hardware Support

[0080] Figure 3 An embodiment of a system utilizing a matrix (tile) operation accelerator is illustrated. In this illustration, host processor / processing system 301 passes command 311 (e.g., matrix manipulation operations such as arithmetic or matrix manipulation operations, or load and store operations) to matrix operation accelerator 307. However, this is shown in this way only for purposes of discussion. As detailed later, accelerator 307 can be part of a processing core. Typically, command 311 as a tile manipulator instruction refers to the tile in register-register (“reg-reg”) or register-memory (“reg-mem”) format. Other commands such as TILESTORE, TILELOAD, TILECONFIG, etc. do not perform data operations on the tile. The command can be a decoded instruction (e.g., micro-operation) or macro-instruction for accelerator 307 to process.

[0081] In this example, coherent memory interface 303 is coupled to host processor / processing system 301 and matrix operation accelerator 307 such that they can share memory. Figure 4 and Figure 5 Different embodiments showing how to use the matrix operation accelerator to share memory are shown. AsFigure 4 As shown, host processor 401 and matrix operation accelerator 405 share the same memory 403. Figure 5 Illustrated is an embodiment in which host processor 501 and matrix operation accelerator 505 do not share memory but can access each other's memory. For example, processor 501 can access on-chip memory 507 and utilize its host memory 503 as normal. Similarly, matrix operation accelerator 505 can access host memory 503, but more typically uses its own memory 507. Note that these memories can be of different types.

[0082] In some embodiments, matrix operation accelerator 307 includes a plurality of FMAs 309 coupled to data buffer 305 (in some implementations, one or more of these buffers 305 are stored in the FMAs of the grid as shown). Data buffer 305 buffers tiles loaded from memory and / or stored to memory (e.g., using on-chip load or on-chip store instructions). The data buffer can be, for example, a plurality of registers. Typically, these FMAs are arranged as a grid of chained FMAs 309 capable of reading and writing tiles. In this example, matrix operation accelerator 307 is used to perform a matrix multiplication operation using tiles T0, T1, and T2. At least one of the tiles is accommodated in FMA grid 309. In some embodiments, all tiles in the operation are stored in FMA grid 309. In other embodiments, only a subset is stored in FMA grid 309. As shown, T1 is accommodated while T0 and T2 are not. Note that A, B, and C refer to the matrices of these tiles, which may or may not occupy the entire space of the tile.

[0083] Figure 6 Illustrated is an embodiment of a matrix multiply-accumulate ("TMMA") operation using tiles.

[0084] The number of rows of the matrix (tile A 601) matches the number of cascaded (chained) FMAs, which includes the latency of the computation. Implementations can freely recycle on a smaller height grid, but the computation remains the same.

[0085] The source / destination vector comes from an N-row tile (tile C 605), and the grid of FMAs 611 performs N vector-matrix operations, resulting in the execution of a complete instruction for matrix multiplication of tiles. Tile B 603 is another vector source and provides "broadcast" terms to the FMAs in each stage.

[0086] In operation, in some embodiments, the elements of matrix B (stored in tile B 603) are scattered across a rectangular grid of FMAs. Matrix B (stored in tile A 601) has its row elements transposed to match the column dimension of the matrix grid of FMAs. At each FMA in the grid, the elements of A and B are multiplied and added to the addend (from above in the figure), and the outgoing sum is passed to the next row of the FMA (or the final output).

[0087] The latency of a single step is proportional to K (the row height of matrix B), and the dependent TMMA typically (within a single die or across dies) has sufficient source-destination rows to hide this latency. The implementation may also split the SIMD (contracted data element) dimension M (the row height of matrix A) across time steps, but this only changes the constant by which K is multiplied. When the program specifies a K smaller than the maximum value enumerated by TMACC, the implementation uses "masking" or "early out" to freely implement this.

[0088] The latency of the entire TMMA is proportional to N*K. The throughput is proportional to N. The number of MACs per TMMA instruction is N*K*M.

[0089] Figure 7 An embodiment is illustrated of an iterative execution of a chained fused multiply-add instruction. Specifically, Figure 7 An iterative execution circuit for one contracted data element location of the destination is illustrated. In this embodiment, the chained fused multiply-add operates on multiple signed sources, where the accumulator is twice the input data size.

[0090] The first signed source (source 1 701) and the second signed source (source 2 703) each have four contracted data elements. Each of these contracted data elements stores signed data such as floating-point data. The third signed source (source 3 709) has two contracted data elements, each of which stores signed data. The sizes of the first signed source 701 and the second signed source 703 are half the size of the third signed source (initial value or previous result) 709. For example, the first signed source 701 and the second signed source 703 may have 32-bit contracted data elements (e.g., single-precision floating-point), while the third signed source 709 may have 64-bit contracted data elements (e.g., double-precision floating-point).

[0091] In this illustration, only the two most significant contracted data element locations of the first signed source 701 and the second signed source 703 and the most significant contracted data element location of the third signed source 709 are shown. Of course, the other contracted data element locations will also be processed.

[0092] As shown, the packed data elements are processed in pairs. For example, a multiplier circuit 705 multiplies the data at the most significant packed data element positions of the first signed source 701 and the second signed source 703, and a multiplier circuit 707 multiplies the data at the next most significant packed data element positions of the first signed source 701 and the second signed source 703. In some embodiments, these multiplier circuits 705 and 707 are reused for other packed data element positions. In other embodiments, additional multiplier circuits are used such that the packed data elements are processed in parallel. In some contexts, a channel sized to the size of the signed third source 709 is used to accomplish the parallel execution. An adder circuit 711 adds the results of each of these multiplications.

[0093] (Using a different adder 713 or the same adder 711), the result of the addition of these multiplication results is added to the data at the most significant packed data element position from the signed source 3 709.

[0094] Finally, the result of the second addition is stored in the signed destination 715 at the packed data element position corresponding to the packed data element position from the signed third source 709, or if there is a next iteration, the result of this second addition is passed on to that next iteration. In some embodiments, a write mask is applied to this storage such that if the corresponding mask (bit) is set, the storage occurs, and if the corresponding mask (bit) is not set, the storage does not occur.

[0095] Figure 8 An embodiment illustrating a subset of the iterative execution of a chained fused multiply-accumulate instruction. Specifically, Figure 8 An execution circuit for an iteration of one packed data element position of a destination is illustrated. In this embodiment, the chained fused multiply-accumulate operates on signed sources, where the accumulator is twice the size of the input data.

[0096] The first signed source (source 1 801) and the second signed source (source 2 803) each have four packed data elements. Each of these packed data elements stores signed data such as integer data. The third signed source (source 3 809) has two packed data elements, each of which stores signed data. The size of the first signed source 801 and the second signed source 803 is half the size of the third signed source 809. For example, the first signed source 801 and the second signed source 803 may have 32-bit packed data elements (e.g., single-precision floating point), while the third signed source 809 may have 64-bit packed data elements (e.g., double-precision floating point).

[0097] In this illustration, only the two most significant packed data element positions of the first signed source 801 and the second signed source 803 and the most significant packed data element position of the third signed source 809 are shown. Of course, other packed data element positions will also be processed.

[0098] As shown, the packed data elements are processed in pairs. For example, the data at the most significant packed data element positions of the first signed source 801 and the second signed source 803 are multiplied using multiplier circuit 805, and the data at the next most significant packed data element positions of the first signed source 801 and the second signed source 803 are multiplied using multiplier circuit 807. In some embodiments, these multiplier circuits 805 and 807 are reused for other packed data element positions. In other embodiments, additional multiplier circuits are used such that the packed data elements are processed in parallel. In some contexts, a channel sized to the size of the signed third source (initial value or previous iteration result) 809 is used to accomplish the parallel execution. The result of each of the multiple multiplications is added to the signed third source 809 using an add / saturate circuit 813.

[0099] When the addition results in an overly large value, the add / saturate (accumulator) circuit 813 preserves the sign of the operands. Specifically, saturation evaluation occurs for the infinite precision result between the multiple additions and the write to the destination or next iteration. When the accumulator 813 is floating point and the input terms are integers, the sum of the products and the floating point accumulator input values are converted to an infinite precision value (a fixed point number of hundreds of bits), the addition of the multiplication result and the third input is performed, and a single rounding to the actual accumulator type is performed.

[0100] Unsigned saturation means that the input value is limited to the largest unsigned number (all 1s) of that element width. Signed saturation means that the value is limited to the range between the smallest complex number and the largest integer of that element width (e.g., for a byte, the range is from -128 (= -2^7) to 127 (= 2^7 - 1)).

[0101] Finally, the result of the addition and saturation check is stored in the signed result 815 in the packed data element position corresponding to the packed data element position of the signed third source 809, or if there is a next iteration, the result is passed on to that next iteration. In some embodiments, a write mask is applied to this storage such that if the corresponding mask (bit) is set, the storage occurs, and if the corresponding mask (bit) is not set, the storage does not occur.

[0102] Figure 9 An embodiment showing a subset of the iterative execution of a chained fused multiply-accumulate instruction. Specifically, Figure 9Execution circuitry for iteratively processing a packed data element position of a depicted destination. In this embodiment, chained fused multiply-accumulates operate on signed and unsigned sources, where the accumulator is four times the size of the input data.

[0103] A first signed source (Source 1 901) and a second unsigned source (Source 2 903) each have four packed data elements. Each of these packed data elements has data such as floating-point data or integer data. A third signed source (initial value or result 915) has a packed data element storing packed data. The size of the first source 901 and the size of the second source 903 are one quarter of the size of the third signed source 915. For example, the first source 901 and the second source 903 may have 16-bit packed data elements (e.g., words), while the third signed source 915 may have 64-bit packed data elements (e.g., double-precision floating point or 64-bit integer).

[0104] In this depiction, only the highest four significant packed data element positions of the first source 901 and the second source 903 and the highest significant packed data element position of the third signed source 915 are shown. Of course, if there are any other packed data element positions, those packed data element positions will also be processed.

[0105] As shown, the packed data elements are processed in quadruples. For example, a multiplier circuit 905 multiplies the data at the highest significant packed data element position of the first source 901 and the second source 903, a multiplier circuit 907 multiplies the data at the next highest significant packed data element position of the first source 901 and the second source 903, a multiplier circuit 909 multiplies the data at the third highest significant packed data element position of the first source 901 and the second source 903, and a multiplier circuit 911 multiplies the data at the lowest significant packed data element position of the first source 901 and the second source 903. In some embodiments, prior to multiplication, the signed packed data elements of the first source 901 are sign-extended and the unsigned packed data elements of the second source 903 are zero-extended.

[0106] In some embodiments, these multiplier circuits 905 - 911 are reused for other packed data element positions. In other embodiments, additional multiplier circuits are used such that the packed data elements are processed in parallel. In some contexts, a channel sized to the size of the signed third source 915 is used to accomplish parallel execution. An adder circuit 911 sums the results of each of these multiplications.

[0107] (Using a different adder 913 or the same adder 911) The result of the addition of these multiplication results is added to the data from the highest significant packed data element position of the signed source 9 915.

[0108] Finally, the result 919 of the second addition is stored into the signed destination at the packed data element position corresponding to the packed data element position of the signed third source 915, or passed to the next iteration. In some embodiments, a write mask is applied to this storage such that if the corresponding mask (bit) is set, the storage occurs, and if the corresponding mask (bit) is not set, the storage does not occur.

[0109] Figure 10 An embodiment illustrating a subset of the execution of an iteration of a chained fused multiply-add instruction. Specifically, Figure 10 An execution circuit for an iteration of a packed data element position of a destination is illustrated. In this embodiment, the chained fused multiply-accumulate operates on signed and unsigned sources, where the accumulator is four times the size of the input data.

[0110] The first signed source (source 1 1001) and the second unsigned source (source 2 1003) each have four packed data elements. Each of these packed data elements stores data such as floating-point data or integer data. The third signed source (initial value or previous result 1015) has a packed data element storing packed data. The size of the first source 1001 and the size of the second source 1003 are one quarter of the size of the third signed source 1015. For example, the first source 1001 and the second source 1003 may have 16-bit packed data elements (e.g., words), while the third signed source 1015 may have 64-bit packed data elements (e.g., double-precision floating-point or 64-bit integer).

[0111] In this illustration, only the four most significant packed data element positions of the first source 1001 and the second source 1003 and the most significant packed data element position of the third signed source 1015 are shown. Of course, if there are any other packed data element positions, these packed data element positions will also be processed.

[0112] As shown, the packed data elements are processed in quadruples. For example, the data at the most significant packed data element position of the first source 1001 and the second source 1003 is multiplied using the multiplier circuit 1005, the data at the next most significant packed data element position of the first source 1001 and the second source 1003 is multiplied using the multiplier circuit 1007, the data at the third most significant packed data element position of the first source 1001 and the second source 1003 is multiplied using the multiplier circuit 1009, and the data at the least significant packed data element position of the first source 1001 and the second source 1003 is multiplied using the multiplier circuit 1011. In some embodiments, before the multiplication, the signed packed data elements of the first source 1001 are sign-extended, and the unsigned packed data elements of the second source 1003 are zero-extended.

[0113] In some embodiments, these multiplier circuits 1005-1011 are reused for other packed data element positions. In other embodiments, additional multiplier circuits are used such that packed data elements are processed in parallel. In some contexts, channels sized to the size of the signed third source 1015 are used to accomplish parallel execution. The result of the addition of these multiplication results is added to the data at the most significant packed data element position from the signed source 3 1015 using the add / saturate circuit 1013.

[0114] When the addition results in a value that is too large or too small for signed saturation, the add / saturate (accumulator) circuit 1013 preserves the sign of the operand. Specifically, saturation evaluation occurs for the infinite precision result between the multiple additions and the write to the destination. When the accumulator 1013 is floating point and the input terms are integer, the sum of the products and the floating point accumulator input values are converted to an infinite precision value (a fixed-point number with hundreds of bits), the addition of the multiplication result and the third input is performed, and a single rounding to the actual accumulator type is performed.

[0115] Finally, the result 1019 of the addition and saturation check is stored into the signed destination in the packed data element position corresponding to the packed data element position of the signed third source 1015, or passed to the next iteration. In some embodiments, a write mask is applied to this storage such that if the corresponding mask (bit) is set, the storage occurs, and if the corresponding mask (bit) is not set, the storage does not occur.

[0116] Figure 11 FIG. illustrates a SIMD implementation of size a power of 2 according to an embodiment, where the accumulator uses an input size larger than the size of the inputs to the multiplier. Note that the (to the multiplier) source and accumulator values can be signed or unsigned values. For an accumulator with a 2X input size (in other words, the size of the accumulator input value is twice the size of the packed data element of the source), Table 1101 illustrates different configurations. For a byte-sized source, the accumulator uses a 16-bit word or a half-precision floating point (HPFP) value. For a word-sized source, the accumulator uses a 32-bit integer or a single-precision floating point (SPFP) value of size 32 bits. For a source of SPFP or 32-bit integer size, the accumulator uses a 64-bit integer or a double-precision floating point (SPFP) value of size 64 bits.

[0117] For an accumulator with a 4X input size (in other words, the size of the accumulator input value is 4 times the size of the packed data elements of the source), Table 1103 illustrates different configurations. For a byte-sized source, the accumulator uses a 32-bit integer or single-precision floating-point (SPFP) value of size 32 bits. For a word-sized source, the accumulator uses a 64-bit integer or double-precision floating-point (DPFP) value of size 64 bits.

[0118] For an accumulator with an 8X input size (in other words, the size of the accumulator input value is 8 times the size of the packed data elements of the source), Table 1105 illustrates the configuration. For a byte-sized source, the accumulator uses a 64-bit integer.

[0119] As previously noted, the matrix operation circuit may be included in the core or may be an external accelerator. Figure 12 An embodiment of a system utilizing a matrix operation circuit is illustrated. In this illustration, multiple entities are coupled to the ring interconnect 1245.

[0120] Multiple cores 1201, 1203, 1205, and 1207 provide non-chip-based instruction support. In some embodiments, the instruction operation circuit 1251 is provided in core 1203, and in other embodiments, the matrix operation circuits 1211 and 1213 are accessible on the ring interconnect 1245.

[0121] In addition, one or more memory controllers 1223 - 1225 are provided to communicate with memories 1233 and 1231 on behalf of the cores and / or matrix operation circuits.

[0122] Figure 13 An embodiment of a processor core pipeline is illustrated, which supports matrix operations using the chip. The branch prediction and decoding circuit 1303 performs branch prediction on instructions stored in the instruction storage 1301, decodes these instructions, and / or performs both branch prediction and decoding. For example, the instructions detailed herein may be stored in the instruction storage. In some implementations, separate circuits are used for branch prediction, and in some embodiments, at least some instructions are decoded into one or more micro-operations, microcode entry points, micro-instructions, other instructions, or other control signals using microcode 1305. The branch prediction and decoding circuit 1303 can be implemented using various different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLAs), microcode read-only memories (ROMs), etc.

[0123] The branch prediction and decoding circuit 1303 is coupled to the rename / allocator circuit 1307, which in some embodiments is coupled to the scheduler circuit 1309. In some embodiments, these circuits provide register renaming, register allocation, and / or scheduling functions by performing one or more of the following steps: 1) renaming logical operand values to physical operand values (e.g., a register alias table in some embodiments); 2) assigning status bits and flags to the decoded instructions; and 3) scheduling the decoded instructions for execution on execution circuits external to the instruction pool (e.g., using reservation stations in some embodiments).

[0124] (The) scheduler circuit(s) 1309 represent any number of different schedulers, including reservation stations, a central instruction window, etc. (The) scheduler unit(s) scheduler circuit(s) 1309 are coupled to (the) physical register file(s) 1315 or include (the) physical register file(s) 1315. Each of (the) physical register file(s) 1315 represents one or more physical register files, where different physical register files store one or more different data types, such as scalar integer, scalar floating point, packed integer, packed floating point, vector integer, vector floating point, status (e.g., an instruction pointer as the address of the next instruction to be executed), slices, etc. In one embodiment, (the) physical register file(s) 1315 include a vector register circuit, a write mask register circuit, and a scalar register circuit. These register circuits may provide architectural vector registers, vector mask registers, and general-purpose registers. (The) physical register file(s) 1315 are covered by the retirement unit 1317 to illustrate various ways in which register renaming and out-of-order execution can be implemented (such as using (the) reorder buffer(s) and (the) retirement register file(s), using (the) future file(s), (the) history buffer(s), (the) retirement register file(s), using register maps and register pools, etc.). The retirement circuit 1317 and (the) physical register file(s) 1315 are coupled to (the) execution circuit(s) 1311.

[0125] Although register renaming has been described in the context of out-of-order execution, it should be understood that register renaming can be used in an in-order architecture. While the illustrated embodiments of the processor may also include separate instruction and data cache units and a shared L2 cache unit, alternative embodiments may also have a single internal cache for both instructions and data, such as, for example, a first-level (L1) internal cache, or multiple levels of internal caches. In some embodiments, the system may include a combination of an internal cache and an external cache external to the core and / or the processor. Alternatively, all caches may be external to the core and / or the processor.

[0126] The execution circuitry 1311 includes one or more sets of execution circuitry 1321, 1323, and 1327 and one or more sets of memory access circuitry 1325. The execution circuitry 1321, 1323, and 1327 perform various operations (e.g., shift, add, subtract, multiply) and operate on various data types (e.g., scalar floating point, packed integer, packed floating point, vector integer, vector floating point). Although some embodiments may include multiple execution units dedicated to a particular function or set of functions, other embodiments may include only one execution unit or multiple execution units that perform all functions. The scalar circuitry 1321 performs scalar operations, the vector / SIMD circuitry 1323 performs vector / SIMD operations, and the matrix operation circuitry 1327 performs matrix (tile) operations as detailed herein.

[0127] As an example, an exemplary register-renamed, out-of-order issue / execution core architecture may implement a pipeline as follows: 1) an instruction fetch circuit performs the fetch and length decode stage; 2) the branch and decode circuit 1303 performs the decode stage; 3) the rename / allocator circuit 1307 performs the allocation stage and the rename stage; 4) the scheduler circuit 1309 performs the schedule stage; 5) the (multiple) physical register files (coupled to or included in the scheduler circuit 1309 and the rename / allocator circuit 1307 and the memory unit) perform the register read / memory read stage; the execution circuitry 1311 performs the execution stage; 6) the memory unit and the (multiple) physical register file units perform the write-back / memory write stage; 7) each unit may be involved in the exception handling stage; and 8) the retirement unit and the (multiple) physical register file units perform the commit stage.

[0128] The core may support one or more instruction sets (e.g., the x86 instruction set (with some extensions added with newer versions); the MIPS instruction set of MIPS Technologies, Inc., Sunnyvale, Calif.; the ARM instruction set of ARM Holdings, Inc., Sunnyvale, Calif. (with optional additional extensions such as NEON)), including the (multiple) instructions described herein. In one embodiment, the core 1390 includes logic for supporting packed data instruction set extensions (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using packed data.

[0129] It should be understood that the core may support multithreading (performing a set of two or more parallel operations or threads), and the multithreading may be accomplished in various ways, including time division multithreading, simultaneous multithreading (where a single physical core provides a logical core for each of the threads for which the physical core is simultaneously multithreading), or a combination thereof (e.g., time division fetching and decoding and thereafter simultaneous multithreading using hyperthreading technology).

[0130] Figure 14 An example of a processor core pipeline is illustrated that supports matrix operations using tiles. The branch prediction and decoding circuit 1403 performs branch prediction on instructions stored in the instruction store 1401, decodes these instructions, and / or performs both branch prediction and decoding. For example, the instructions detailed herein may be stored in the instruction store. In some implementations, separate circuitry is used for branch prediction, and in some embodiments, at least some of the instructions are decoded into one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals using microcode 1405. The branch prediction and decoding circuit 1403 may be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLAs), microcode read-only memories (ROMs), and the like.

[0131] The branch prediction and decoding circuit 1403 is coupled to the rename / allocator circuit 1407, which in some embodiments is coupled to the scheduler circuit 1409. In some embodiments, these circuits provide register renaming, register allocation, and / or scheduling functions by performing one or more of the following steps: 1) renaming logical operand values to physical operand values (e.g., a register alias table in some embodiments); 2) assigning status bits and flags to the decoded instructions; and 3) scheduling the decoded instructions for execution on execution circuitry external to the instruction pool (e.g., using reservation stations in some embodiments).

[0132] (Multiple) scheduler circuits 1409 represent any number of different schedulers, including reservation stations, central instruction windows, etc. (Multiple) scheduler unit scheduler circuits 1409 are coupled to (multiple) physical register files 1415 or include (multiple) physical register files 1415. Each of (multiple) physical register files 1415 represents one or more physical register files, where different physical register files store one or more different data types, such as scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points, status (e.g., instruction pointer as the address of the next instruction to be executed), slices, etc. In one embodiment, (multiple) physical register files 1415 include vector register circuits, write mask register circuits, and scalar register circuits. These register circuits can provide architectural vector registers, vector mask registers, and general-purpose registers. (Multiple) physical register files 1415 are covered by a retirement unit 1417 to illustrate various ways to implement register renaming and out-of-order execution (such as using (multiple) reorder buffers and (multiple) retirement register files, using (multiple) future files, (multiple) history buffers, (multiple) retirement register files, using register maps and register pools, etc.). Retirement circuit 1417 and (multiple) physical register files 1415 are coupled to (multiple) execution circuits 1411.

[0133] Although register renaming is described in the context of out-of-order execution, it should be understood that register renaming can be used in an in-order architecture. While the illustrated embodiments of the processor may also include separate instruction and data cache units and a shared L2 cache unit, alternative embodiments may also have a single internal cache for both instructions and data, such as, for example, a first-level (L1) internal cache, or multiple levels of internal caches. In some embodiments, the system may include a combination of an internal cache and an external cache outside the core and / or the processor. Alternatively, all caches may be outside the core and / or the processor.

[0134] Execution circuits 1411 include a set of one or more execution circuits 1427 and a set of one or more memory access circuits 1325. Execution circuits 1427 perform the vector (slice) operations detailed herein.

[0135] As an example, an exemplary register-renamed, out-of-order issue / execution core architecture can implement a pipeline as follows: 1) An instruction fetch circuit performs a fetch and length decode stage; 2) A branch and decode circuit 1403 performs a decode stage; 3) A rename / allocator circuit 1407 performs an allocation stage and a rename stage; 4) A scheduler circuit 1409 performs a schedule stage; 5) (A plurality of) physical register files (coupled to or included in the scheduler circuit 1409 and the rename / allocator circuit 1407 and the memory unit) perform a register read / memory read stage; An execution circuit 1411 performs an execution stage; 6) A memory unit and (a plurality of) physical register file units perform a write-back / memory write stage; 7) Each unit may be involved in an exception handling stage; and 8) A retirement unit and (a plurality of) physical register file units perform a commit stage.

[0136] The core may support one or more instruction sets (e.g., the x86 instruction set (with some extensions added with newer versions); the MIPS instruction set of MIPS Technologies, Inc., in Sunnyvale, California; the ARM instruction set of ARM Holdings, Inc., in Sunnyvale, California (with optional additional extensions such as NEON)), including the (multiple) instructions described herein. In one embodiment, the core 1490 includes logic for supporting a packed data instruction set extension (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using packed data.

[0137] It should be understood that the core may support multithreading (executing a collection of two or more parallel operations or threads), and this multithreading can be accomplished in various ways, including time division multithreading, simultaneous multithreading (where a single physical core provides a logical core for each of the threads for which the physical core is simultaneously multithreading), or a combination thereof (e.g., time division fetching and decoding and thereafter simultaneous multithreading using hyperthreading technology).

[0138] B. Layout

[0139] Throughout this specification, a row-major data layout is used to represent data. Column-major users should transform the items according to the orientation of the items. Figure 15 An example of a matrix expressed in row-major format and column-major format is illustrated. As shown, matrix A is a 2x3 matrix. When the matrix is stored in row-major format, the data elements of the rows are contiguous. When the matrix is stored in column-major format, the data elements of the columns are contiguous. A T *B T =(BA) T is a well-known property of matrices, where the superscript T represents transpose. Reading column-major data as if it were row-major data results in a matrix that looks like the transpose matrix.

[0140] In some embodiments, the behavior-major semantics are utilized in hardware, and data that is column-major will swap the operand order and make the result the transpose of the matrix, but for subsequent column-major reads from memory, it is the correct non-transposed matrix.

[0141] For example, if there are two column-major matrices to be multiplied:

[0142]

[0143] The data matrix will be stored in linear memory as follows (column-major):

[0144] a c e b d f

[0145] And

[0146] g h i j k l.

[0147] Reading those matrices as column-major with dimensions 2x3 and 3x2, they appear as:

[0148] a c e and g h

[0149] b d f i j

[0150] k l

[0151] Swapping the order and matrix multiplication:

[0152] g h a c e ag+bh cg+dh eg+fh

[0153] i j * b d f = ai+bj ci+dj ei+fj

[0154] k l ak+bl ck+dl ek+fl

[0155] The transposed matrix is taken out and can then be stored in row-major order:

[0156] ag+bh cg+dh eg+fh ai+bj ci+dj ei+fj ak+bl ck+dl ek+fl

[0157] And is used in subsequent column-major calculations, which is the correct non-transposed matrix:

[0158] ag+bh ai+bj ak+bl

[0159] cg+dh ci+dj ck+dl

[0160] eg+fh ei+fj ek+fl

[0161] III. Exemplary Use

[0162] Figure 16 An example of the use of the illustrated matrix (tile) is shown. In this example, matrix C 1601 includes two tiles, matrix A 1603 includes one tile, and matrix B 1605 includes two tiles. This illustration shows an example of the inner loop of an algorithm for calculating matrix multiplication. In this example, two result tiles tmm0 and tmm1 from matrix C 1601 are used to accumulate intermediate results. When one tile (tmm2) from matrix B 1603 is multiplied by two tiles from matrix B 1605, this tile is reused 2 times. Pointers are used to load new A tiles and two new B tiles from the directions indicated by the arrows. The outer loop adjustment for the C tiles, which is not shown, is used to adjust the pointers.

[0163] The exemplary code shown in the figure includes the use of tile configuration instructions and is executed to configure tile usage, load tiles, loop for processing tiles, store tiles into memory, and release tile usage.

[0164] Figure 17 An embodiment of the use of the illustrated matrix (tile) is shown. At 1701, tile usage is configured. For example, the TILECONFIG instruction is executed to configure tile usage, including setting the number of rows and columns of each tile. Typically, at 1703, at least one matrix (tile) is loaded from memory. At 1705, at least one matrix (tile) operation is performed using the matrix (tile). At 1707, at least one matrix (tile) is stored out to memory, and at 1709, a context switch may occur.

[0165] IV. Exemplary Configuration

[0166] A. Tile Configuration Hardware Support

[0167] As discussed above, tile usage typically requires configuration before use. For example, it may not be necessary to use all rows and columns fully. Not configuring these rows and columns in some embodiments not only saves power but also allows the configuration to determine whether an operation will generate an error. For example, matrix multiplication in the form of (NxM)*(L*N) will generally not work if M and L are not the same.

[0168] Before using a matrix of tiles, in some embodiments, tile support will be configured. For example, configure how many rows and columns each tile will have, which tiles will be used, and so on. The TILECONFIG instruction is an improvement to the computer itself because it provides support for configuring the computer to use a matrix accelerator (either as part of a processor core or as an external device). Specifically, the execution of the TILECONFIG instruction causes the configuration to be fetched from memory and applied to the matrix (tile) settings within the matrix accelerator.

[0169] i. Tile usage configuration

[0170] Figure 18 Illustrates support for the configuration of tile usage according to an embodiment. Memory 1801 contains a description 1803 of the matrix (tiles) to be supported.

[0171] The execution circuitry 1811 of the processor / core 1805 stores multiple aspects of the tile description 1803 into the tile configuration 1817. The tile configuration 1817 details which tiles are configured for the palette (the number of rows and columns in each tile) and the flag for matrix support in use. Specifically, the instruction execution resource 1811 is configured to use the tiles as specified by the tile configuration 1817. The instruction execution resource may also include machine-specific registers or configuration registers for indicating tile usage. Additional values are also set, such as the in-use value and the start value. The tile configuration 1817 uses one or more registers 1819 to store tile usage and configuration information.

[0172] Figure 19 Illustrates an embodiment of the description of the matrix (tiles) to be supported. This is the description that will be stored in response to the execution of the STTILECFG instruction. In this example, each field is a byte. In byte [0], the palette ID 1901 is stored. The palette ID is used to index into the palette table 1813, which stores the number of bytes in the tile and the number of bytes per row of the tile associated with that ID as defined by the configuration.

[0173] Byte 1 stores the value to be stored in the "startRow" register 1903, and byte 2 stores the value to be stored in the "startP" register 1905. To support resuming instructions after these events, these instructions store information in these registers. To support resuming instructions after interrupt events such as those detailed above, these instructions store information in these registers. The startRow value indicates the row that should be used for resuming. The startP value indicates the position within the row used for storage operations when in use, and in some embodiments, the startP value indicates the lower half of the row (in the lower slice of the pair) or the upper half of the row (in the higher slice of the pair). Generally, this position within the row (column) is not required.

[0174] Successfully executing a matrix (tile) instruction will set both startRow and startP to zero, with the exceptions of TILECONFIG and STTILECFG.

[0175] At any time when not resuming an interrupted matrix (tile), it is the responsibility of the software to zero out the startRow and startP values. For example, an unmasked floating-point exception handler may decide to complete the operation in software and change the program counter value to another instruction, typically the next instruction. In this case, before resuming the program, the software exception handler must zero out the startRow and startP values in the exception presented to the software exception handler by the operating system. The operating system will then use the resume instruction to reload those values.

[0176] Byte 3 stores the indication of the pair of tiles (1b per tile) 1907.

[0177] Bytes 16 - 17 store the number of rows 1913 and columns 1915 of tile 0, bytes 18 - 19 store the number of rows and columns of tile 1, and so on. In other words, each 2-byte group specifies the number of rows and columns of a tile. If a 2-byte group is not used to specify a parameter, they should have a value of zero. Specifying tile parameters for more tiles than the implementation limit or palette limit results in an error. Unconfigured tiles are set to the initial state with 0 rows and 0 columns.

[0178] Finally, the configuration in memory typically ends with an end-of-description such as all zeros for several consecutive bytes.

[0179] ii. Exemplary Tile and Tile Configuration Storage

[0180] Figure 20(A)-Figure 20(D)Illustration of an example of registers 1819. FIG. 20(A) illustrates multiple registers 1819. As shown, each tile (TMM0 2001... TMMN 2003) has separate registers, where each register stores the row size and column size of that particular tile. StartP and StartRow are stored in separate registers 2011 and 2013. One or more status registers 2015 are set (e.g., TILES_CONFIGURED = 1) to indicate that the tiles are configured for use.

[0181] FIG. 20(B) illustrates multiple registers 1819. As shown, each tile has separate registers for its rows and its columns. For example, the TMM0 row configuration 2021, the TMM0 column configuration 2023, StartP, and StartRow are stored in separate registers 2011 and 2013. One or more status registers 2015 are set (e.g., TILES_CONFIGURED = 1) to indicate that the tiles are configured for use.

[0182] FIG. 20(C) illustrates a single register 1819. As shown, the register stores the tile configuration (rows and columns per tile) 2031, StartP 2011, and StartRow 2013 in a single register as a packed data register. One or more status registers 2015 are set (e.g., TILES_CONFIGURED = 1) to indicate that the tiles are configured for use.

[0183] FIG. 20(D) illustrates multiple registers 1819. As shown, a single register stores the tile configuration (rows and columns per tile) 2031. StartP and StartRow are stored in separate registers 2011 and 2013. One or more status registers 2015 are set (e.g., TILES_CONFIGURED = 1) to indicate that the tiles are configured for use.

[0184] Other combinations are envisioned, such as combining the start registers into a single register in which the start registers are shown separately, and so on.

[0185] Exemplary execution

[0186] Figure 21Illustrated is an exemplary execution of the TZPAIR instruction. The TZPAIR instruction format includes fields for an opcode and a destination tile identifier tmm0 that is used to identify a pair of tiles having M rows, N columns, and both PAIR and VALID parameters set to true. As shown, execution circuit 2104 receives the decoded TZPAIR instruction 2102 and, in some embodiments, execution circuit 2104 uses a grid 2106 of FMAs to write zeros to each element of a left destination matrix (tile) 2108 and a right destination matrix (tile) 2110. As detailed above, the left destination matrix (tile) and the right destination matrix (tile) may be stored in a set of registers, a location in memory, or other storage accessible to the execution circuit.

[0187] As shown, execution circuit 2104 executes the decoded TZPAIR instruction 2102 to zero the elements of the left destination matrix (tile) 2108 and the right destination matrix (tile) 2110.

[0188] Also shown is that the remaining (unconfigured) columns and rows are set to zero, which is done in some embodiments. In some embodiments, a matrix (tile) is configured to use only a subset of the possible rows and columns. For example, a matrix (tile) may have up to 16 rows and 16 columns available for use, but only 4 rows and 4 columns of each matrix (tile) are used. Configuration of each matrix (tile) is typically done by executing a configuration instruction prior to use of the matrix (tile). In this example, there are N possible columns and M possible rows.

[0189] II. Exemplary Instruction Formats

[0190] An example of the format of the TZPAIR instruction is TZPAIR{B / W / D / Q}TMM0. In some embodiments, TZPAIR{B / W / D / Q} is an opcode mnemonic for the instruction, where B / W / D / Q indicates the data element size (byte, word, double word, quad word) of the two destinations. In some embodiments, the TMM2 field is an R / M value (such as Figure 25A-Figure 25B 2546), the TMM1 field is Figure 25A-Figure 25B REG 2544 of Figure 25A-Figure 25B and the data element size is found in

[0191] Figure 25A-Figure 25B Figure 25A-Figure 25Bfield 2550). In one embodiment, an SIB-type memory operand may include an encoding that identifies a base register. The content of the base register may represent a base address in memory, and the address of a specific destination location in memory may be calculated through this base address. For example, the base address may be the address of the first location in a block of potential destination locations for an extended vector instruction. In one embodiment, an SIB-type memory operand may include an encoding that identifies an index register. Each element of the index register may specify an index or offset value that can be used to calculate the address of the corresponding destination address within the block of potential destination addresses through the base address. In one embodiment, an SIB-type memory operand may include an encoding that specifies a scale factor to be used for each index value when calculating the corresponding destination address. For example, if a scale factor value of 4 is encoded in the SIB-type memory operand, each index value obtained from an element of the index register is multiplied by 4 and then added to the base address to calculate the destination address.

[0192] In one embodiment, an SIB-type memory operand of the form vm32{x,y,z} may identify an array of vector memory operands specified using SIB-type memory addressing. In this example, an array of memory addresses is specified using a common base register, a constant scale factor, and a vector index register with elements each containing an index value that is 32 bits in length. The vector index register may be a 128-bit (e.g., XMM) register (vm32x), a 256-bit (e.g., YMM) register (vm32y), or a 512-bit (e.g., ZMM) register (vm32z). In another embodiment, an SIB-type memory operand of the form vm64{x,y,z} may identify an array of vector memory operands specified using SIB-type memory addressing. In this example, an array of memory addresses is specified using a common base register, a constant scale factor, and a vector index register with elements each containing an index value that is 64 bits in length. The vector index register may be a 128-bit (e.g., XMM) register (vm64x), a 256-bit (e.g., YMM) register (vm64y), or a 512-bit (e.g., ZMM) register (vm64z).

[0193] III. Exemplary Method(s) Executed

[0194] Figure 22 The figure illustrates an embodiment of a method for processing a TZPAIR instruction executed by a processor.

[0195] At 2201, an instruction is fetched. For example, a TZPAIR instruction is fetched, which has fields for an opcode and a destination matrix (slice) operand, and the destination matrix (slice) operand has a pair parameter equal to true. In some embodiments, the instruction is fetched from an instruction cache. The opcode of the TZPAIR instruction indicates to zero the packed data element positions of the left matrix (slice) and the right matrix (slice) of the identified destination matrix (slice). In some embodiments, the opcode of the TZPAIR instruction includes a suffix, such as "B", "W", "D", or "Q", to specify the size of each slice element as a byte, word, double word, and quad word, respectively.

[0196] At 2203, the fetched instruction is decoded. For example, the fetched TZPAIR instruction is decoded by a decoding circuit such as those detailed herein.

[0197] At 2205, the execution of the decoded instruction is scheduled (as needed).

[0198] At 2207, the decoded TZPAIR instruction is executed by an execution circuit (hardware) such as those detailed herein. For the TZPAIR instruction, this execution will cause the execution circuit to zero the elements of the left matrix (slice) and the right matrix (slice) of the identified target matrix (slice). In some embodiments, the unconfigured elements of the rows of the destination matrix (slice) are also zeroed.

[0199] In some embodiments, at 2209, the instruction is committed or retired.

[0200] Figure 23 A more detailed description of the execution of the TZPAIR instruction is illustrated. Generally, this is performed by an execution circuit such as those detailed above.

[0201] At 2302, it is determined whether all of the following conditions are true: 1) Is there at least one configured matrix (slice)? 2) Does the identified destination matrix (slice) have a VALID parameter set to TRUE? And 3) Does the identified destination matrix (slice) have a PAIR parameter set to TRUE? When any of these conditions is not true, an error is generated at 2304.

[0202] When all conditions tested at 2302 are true, the execution circuit loops on each row M of the left matrix (slice) and the right matrix (slice) of the identified destination matrix (slice) starting from the first row at 2306. For each row, the execution circuit executes an inner loop at 2308, starting from the first column, and loops on each column N of the left matrix (slice) and the right matrix (slice) of the identified destination matrix (slice). For each element of the inner loop, the execution circuit determines at 2310 how many bytes are contained in each element of the destination matrix (slice). For example, the TZPAIR instruction may include an operand, an opcode prefix, or an opcode suffix that is one of "B", "W", "D", and "Q", which are used to specify element sizes of 1 byte, 2 bytes, 4 bytes, and 8 bytes, respectively. When the left matrix (slice) elements and the right matrix (slice) elements of the identified destination matrix (slice) contain 1 byte, the execution circuit sets each byte-sized element to zero at 2312. When the left matrix (slice) elements and the right matrix (slice) elements of the identified destination matrix (slice) contain 2 bytes, the execution circuit sets each word-sized element to zero at 2314. When the left matrix (slice) elements and the right matrix (slice) elements of the identified destination matrix (slice) contain 4 bytes, the execution circuit sets each double-word-sized element to zero at 2316. When the left matrix (slice) elements and the right matrix (slice) elements of the identified destination matrix (slice) contain 8 bytes, the execution circuit sets each quad-word-sized element to zero at 2318.

[0203] After setting the elements of the left matrix (slice) and the right matrix (slice) of the identified destination matrix (slice) at one of 2312, 2314, 2316, and 2318, the execution circuit increments N at 2320 and determines whether there are any columns remaining in the inner loop. If so, it continues to 2308 to execute the next iteration of the inner loop. However, when the determination at 2320 indicates that there are no columns remaining, the execution circuit increments M at 2322 and determines whether there are any rows remaining in the outer loop, and if so, it continues to 2306 to execute the next iteration of the outer loop. However, when the determination at 2322 indicates that there are no rows remaining, the process ends.

[0204] IV. Exemplary Pseudocode

[0205] Figure 24It is exemplary pseudocode describing an embodiment of a method for processing TZPAIR instructions executed by a processor. As shown in pseudocode 2402, the TZPAIR instruction includes an opcode TZPAIRB and a destination matrix (slice) identifier for identifying a configured matrix (slice) having VALID and PAIR parameters set to TRUE. The suffix "B" included in the opcode indicates that the destination matrix (slice) includes a left matrix (slice) tmm1.left and a right matrix (slice) tmm1.right with elements of byte size. As shown, if any of the three error checks fails, pseudocode 2402 first causes the execution circuit to generate an error. Subsequently, the pseudocode causes the processor to loop over each row j and each column k of the left matrix (slice) tmm1.left and the right matrix (slice) tmm1.right. At each element, the processor sets the elements at tmm1.left[j][k] and tmm1.right[j][k] to zero.

[0206] Pseudocode 2404 operates similarly to pseudocode 2402, but processes an instruction with an opcode having a "W" suffix, which indicates that the elements of the left matrix (slice) tmm1.left and the right matrix (slice) tmm1.right are each two bytes in size.

[0207] Pseudocode 2406 operates similarly to pseudocode 2402, but processes an instruction with an opcode having a "D" suffix, which indicates that the elements of the left matrix (slice) tmm1.left and the right matrix (slice) tmm1.right are each four bytes in size.

[0208] Pseudocode 2408 operates similarly to pseudocode 2402, but processes an instruction with an opcode having a "Q" suffix, which indicates that the elements of the left matrix (slice) tmm1.left and the right matrix (slice) tmm1.right are each eight bytes in size. Pseudocode 2408 performs zeroing in a slightly different order than the order of pseudocode 2402, as the first loop zeros all elements of the left matrix (slice) and the second loop zeros all elements of the right matrix (slice).

[0209] Further examples

[0210] Example 1 provides a processor including: a decoding circuit for decoding a matrix pair zeroing instruction having fields for an opcode and an identifier for identifying a destination matrix, the destination matrix for having a PAIR parameter equal to true; and an execution circuit for executing the decoded matrix pair zeroing instruction to zero each element of the left and right matrices of the identified destination matrix.

[0211] Example 2 includes an entity of the exemplary processor of Example 1, wherein the opcode defines the size of each data element of the left matrix and the right matrix.

[0212] Example 3 includes an entity of the exemplary processor of Example 2, wherein the size of each data element of the left matrix and the right matrix is a double word.

[0213] Example 4 includes an entity of the exemplary processor of Example 2, wherein the size of each data element of the left matrix and the right matrix is a word.

[0214] Example 5 includes an entity of the exemplary processor of any one of Examples 1-4, wherein the execution circuit is further configured to zero any data elements in the remaining columns of the left matrix and the right matrix and in the unconfigured rows of the left matrix and the right matrix.

[0215] Example 6 includes an entity of the exemplary processor of any one of Examples 1-4, wherein each of the left matrix and the right matrix is a plurality of registers for representing a matrix.

[0216] Example 7 includes an entity of the exemplary processor of any one of Examples 1-4, wherein the execution circuit is configured to generate an error when at least one of the following is determined: the number of configured slices is equal to zero; the PAIR parameter of the identified destination matrix is not set to true; and the VALID parameter of the identified destination matrix is not set to true.

[0217] Example 8 provides a method, including: decoding a matrix pair zeroing instruction having fields for an opcode and a destination matrix identifier for identifying a destination matrix having a PAIR parameter equal to true; and executing the decoded matrix pair zeroing instruction to zero each element of the left matrix and the right matrix of the identified destination matrix.

[0218] Example 9 includes an entity of the exemplary method of Example 8, wherein the opcode defines the size of each data element of the left matrix and the right matrix.

[0219] Example 10 includes an entity of the exemplary method of Example 9, wherein the size of each data element of the left matrix and the right matrix is a double word.

[0220] Example 11 includes an entity of the exemplary method of Example 9, wherein the size of each data element of the left matrix and the right matrix is a word.

[0221] Example 12 includes an entity of the exemplary method of any one of Examples 8-11, further including zeroing any data elements in the remaining columns of the left matrix and the right matrix and in the unconfigured rows of the left matrix and the right matrix.

[0222] Example 13 includes an entity of the exemplary method of any one of Examples 8-11, wherein each of the left matrix and the right matrix is a plurality of registers for representing a matrix.

[0223] Example 14 includes an entity of the exemplary method of any one of Examples 8-11, and further includes: generating an error when at least one of the following is determined: the number of configured slices is equal to zero; the PAIR parameter of the identified destination matrix is not set to true; and the VALID parameter of the identified destination matrix is not set to true.

[0224] Example 15 provides a non-transitory machine-readable medium that stores a matrix pair zeroing instruction, and the matrix pair zeroing instruction causes a processor to execute the instruction by the following steps: decoding the matrix pair zeroing instruction, the matrix pair zeroing instruction having fields for an opcode and a destination matrix identifier, the destination matrix identifier being used to identify a destination matrix having a PAIR parameter equal to true; and executing the decoded matrix pair zeroing instruction to zero each element of the left matrix and the right matrix of the identified destination matrix.

[0225] Example 16 includes an entity of the exemplary non-transitory machine-readable medium of Example 15, wherein the opcode defines the size of each data element of the left matrix and the right matrix.

[0226] Example 17 includes an entity of the exemplary non-transitory machine-readable medium of Example 15, wherein the size of each data element of the left matrix and the right matrix is a double word.

[0227] Example 18 includes an entity of the exemplary non-transitory machine-readable medium of Example 15, wherein the size of each data element of the left matrix and the right matrix is a word.

[0228] Example 19 includes an entity of the exemplary non-transitory machine-readable medium of any one of Examples 15-18, and further includes zeroing the remaining columns of the left matrix and the right matrix and any data elements in the unconfigured rows of the left matrix and the right matrix.

[0229] Example 20 includes an entity of the exemplary non-transitory machine-readable medium of any one of Examples 15-18, wherein each of the left matrix and the right matrix is a plurality of registers for representing a matrix.

[0230] Example 21 includes an entity of the exemplary non-transitory machine-readable medium of any one of Examples 15-18, and further includes: generating an error when at least one of the following is determined: the number of configured slices is equal to zero; the PAIR parameter of the identified destination matrix is not set to true; and the VALID parameter of the identified destination matrix is not set to true.

[0231] Example 22 provides a system including: a processor and an accelerator coupled to the processor, the accelerator including: a decoding circuit for decoding a matrix pair zeroing instruction having fields for an opcode and a destination matrix identifier for identifying a destination matrix having a PAIR parameter equal to true; and an execution circuit for executing the decoded matrix pair zeroing instruction to zero each element of the left and right matrices of the identified destination matrix.

[0232] Example 23 includes an entity of the exemplary system of Example 22, wherein the opcode defines the size of each data element of the left and right matrices.

[0233] Example 24 includes an entity of the exemplary system of any one of Examples 22-23, wherein the execution circuit is further configured to zero any data elements in the remaining columns of the left and right matrices and in the unconfigured rows of the left and right matrices.

[0234] Example 25 includes an entity of the exemplary system of any one of Examples 22-23, wherein the left and right matrices are each a plurality of registers for representing a matrix.

[0235] V. Detailed Exemplary Systems, Processors, and Simulations

[0236] Examples of hardware, software, etc. for performing the instructions described above are detailed herein. For example, the following description details various aspects of instruction execution, including various pipeline stages such as fetch, decode, schedule, execute, retire, etc.

[0237] Instruction Set

[0238] An instruction set can include one or more instruction formats. A given instruction format can define various fields (e.g., number of bits, position of bits) to specify the operation to be performed (e.g., opcode) and the operand(s) and / or other data field(s) (e.g., mask) on which the operation is to be performed, etc. Some instruction formats are further decomposed by the definition of instruction templates (or sub-formats). For example, an instruction template of a given instruction format can be defined as a different subset of the fields of that instruction format (the included fields are generally in the same order, but at least some fields have different bit positions because fewer fields are included), and / or defined as having a given field interpreted in a different way. Thus, each instruction of the ISA is expressed using a given instruction format (and if defined, according to a given one of the instruction templates in that instruction format), and includes fields for specifying the operation and operands. For example, an exemplary ADD (addition) instruction has a specific opcode and instruction format, and the specific instruction format includes an opcode field for specifying the opcode and operand fields for selecting the operands (source 1 / destination and source 2); and the appearance of the ADD instruction in the instruction stream will cause specific content for selecting specific operands to be in the operand fields. There have been introduced and / or published SIMD extension sets known as Advanced Vector Extensions (AVX) (AVX1 and AVX2) and using the Vector Extension (VEX) encoding scheme (see, for example, the 64 and IA-32 Architectures Software Developer's Manual in September 2014; and see the Advanced Vector Extensions Programming Reference in October 2014).

[0239] Exemplary instruction format

[0240] Embodiments of the instruction(s) described herein can be embodied in different formats. Additionally, exemplary systems, architectures, and pipelines are described in detail below. Embodiments of the instruction(s) can be executed on such systems, architectures, and pipelines, but are not limited to those detailed.

[0241] General vector-friendly instruction format

[0242] A vector-friendly instruction format is an instruction format suitable for vector instructions (e.g., there are specific fields dedicated to vector operations). Although embodiments are described in which both vector and scalar operations are supported by the vector-friendly instruction format, alternative embodiments use only vector operations via the vector-friendly instruction format.

[0243] Figure 25A-25B is a block diagram showing a general vector-friendly instruction format and its instruction templates according to an embodiment of the present invention. Figure 25Ais a block diagram showing a general vector friendly instruction format and its Class A instruction templates according to an embodiment of the present invention; and Figure 25B is a block diagram showing a general vector friendly instruction format and its Class B instruction templates according to an embodiment of the present invention. Specifically, Class A and Class B instruction templates are defined for the general vector friendly instruction format 2500, both of which include instruction templates for memoryless access 2505 and instruction templates for memory access 2520. The term "general" in the context of the vector friendly instruction format refers to an instruction format that is not tied to any specific instruction set.

[0244] Although embodiments of the present invention will be described in which the vector friendly instruction format supports the following: a 64-byte vector operand length (or size) with a 32-bit (4-byte) or 64-bit (8-byte) data element width (or size) (and thus, a 64-byte vector consists of 16 double-word sized elements, or alternatively, consists of 8 quad-word sized elements); a 64-byte vector operand length (or size) with a 16-bit (2-byte) or 8-bit (1-byte) data element width (or size); a 32-byte vector operand length (or size) with a 32-bit (4-byte), 64-bit (8-byte), 16-bit (2-byte), or 8-bit (1-byte) data element width (or size); and a 16-byte vector operand length (or size) with a 32-bit (4-byte), 64-bit (8-byte), 16-bit (2-byte), or 8-bit (1-byte) data element width (or size); however, alternative embodiments may support larger, smaller, and / or different vector operand sizes (e.g., 256-byte vector operands) with larger, smaller, or different data element widths (e.g., a 128-bit (16-byte) data element width).

[0245] Figure 25A The Class A instruction templates in include: 1) within the instruction template for memoryless access 2505, an instruction template showing a fully rounded control type operation 2510 without memory access and an instruction template for a data transformation type operation 2515 without memory access; and 2) within the instruction template for memory access 2520, an instruction template showing the timeliness 2525 of memory access and an instruction template for the non-timeliness 2530 of memory access. Figure 25B The Class B instruction templates in include: 1) within the instruction template for memoryless access 2505, an instruction template showing a partially rounded control type operation 2512 with write mask control without memory access and an instruction template for a vsize (vector size) type operation 2517 with write mask control without memory access; and 2) within the instruction template for memory access 2520, an instruction template showing the write mask control 2527 of memory access.

[0246] The general vector friendly instruction format 2500 includes the following as described belowFigure 25A-Figure 25B The following fields listed in the order illustrated therein.

[0247] Format field 2540 - The specific value (instruction format identifier value) in this field uniquely identifies the vector-friendly instruction format, and thereby identifies that the instruction appears in the instruction stream in the vector-friendly instruction format. Thus, this field is optional in the sense that it is not required for an instruction set that only has a general vector-friendly instruction format.

[0248] Base operation field 2542 - Its content differentiates different base operations.

[0249] Register index field 2544 – Its content directly or through address generation specifies the location of source or destination operands in registers or in memory. These fields include a sufficient number of bits to select N registers from a PxQ (e.g., 32x512, 16x128, 32x1024, 64x1024) register file. Although in one embodiment N can be up to three source registers and one destination register, alternative embodiments can support more or fewer source and destination registers (e.g., can support up to two source registers, where one of these source registers also serves as the destination register; can support up to three source registers, where one of these source registers also serves as the destination register; can support up to two source registers and one destination register).

[0250] Modifier field 2546 - Its content differentiates instructions that appear in the general vector instruction format and specify memory access from those that do not; i.e., it differentiates between instruction templates without memory access 2505 and instruction templates with memory access 2520. Memory access operations read and / or write to the memory hierarchy (in some cases, using values in registers to specify source and / or destination addresses), while non-memory access operations do not (e.g., the source and / or destination are registers). Although in one embodiment, this field also selects between three different ways to perform memory address calculation, alternative embodiments can support more, fewer, or different ways to perform memory address calculation.

[0251] Extended operation field 2550 - Its content differentiates which one of various different operations to perform in addition to the base operation. This field is context-dependent. In one embodiment of the present invention, this field is divided into a class field 2568, an α field 2552, and a β field 2554. The extended operation field 2550 allows multiple sets of common operations to be performed in a single instruction rather than in 2, 3, or 4 instructions.

[0252] Scale field 2560 - The content thereof is allowed to scale the content of the index field used for memory address generation (e.g., for address generation using 2 比例 * index + base address).

[0253] Displacement field 2562A - The content thereof is used as part of memory address generation (e.g., for address generation using 2 比例 * index + base address + displacement).

[0254] Displacement factor field 2562B (note that the juxtaposition of displacement field 2562A directly on displacement factor field 2562B indicates the use of one or the other) - The content thereof is used as part of address generation; it specifies a displacement factor scaled by the size (N) of the memory access, where N is the number of bytes in the memory access (e.g., for address generation using 2 比例 * index + base address + scaled displacement). Redundant low-order bits are ignored, and thus the content of the displacement factor field is multiplied by the total size (N) of the memory operand to generate the final displacement to be used in calculating the effective address. The value of N is determined by the processor hardware at runtime based on the full opcode field 2574 (described later in this document) and the data manipulation field 2554C. The displacement field 2562A and the displacement factor field 2562B are optional in the sense that they are not used in instruction templates without memory access 2505 and / or different embodiments may implement only one of the two or neither of the two.

[0255] Data element width field 2564 - The content thereof differentiates which of multiple data element widths is used (in some embodiments for all instructions; in other embodiments only for some instructions). This field is optional in the sense that it is not needed if only one data element width is supported and / or some aspect of the opcode is used to support the data element width.

[0256] Write mask field 2570 - whose content controls whether the data element positions in the destination vector operand reflect the results of the base operation and the expansion operation on a per data element position basis. Class A instruction templates support merge-write masking, while Class B instruction templates support both merge-write masking and zero-write masking. When merging, the vector mask allows any set of elements in the destination to be protected from update during the execution of any operation (specified by the base operation and the expansion operation); in another embodiment, the old value of each element of the destination where the corresponding mask bit has 0 is maintained. In contrast, when zeroing, the vector mask allows any set of elements in the destination to be zeroed during the execution of any operation (specified by the base operation and the expansion operation); in one embodiment, the elements of the destination are set to 0 when the corresponding mask bit has a value of 0. A subset of this functionality is the ability to control the vector length of the operation being performed (i.e., the span from the first to the last element being modified); however, the elements being modified do not have to be contiguous. Thus, the write mask field 2570 allows partial vector operations, which include loads, stores, arithmetic, logic, etc. Although embodiments of the invention have been described in which the content of the write mask field 2570 selects one of a plurality of write mask registers containing the write mask to be used (and thus, the content of the write mask field 2570 indirectly identifies the masking to be performed), alternative embodiments alternatively or additionally allow the content of the mask write field 2570 to directly specify the masking to be performed.

[0257] Immediate number field 2572 - whose content allows the specification of an immediate number. This field is optional in the sense that it does not exist in a general vector friendly format that does not support immediate numbers and does not exist in instructions that do not use immediate numbers.

[0258] Class field 2568 - whose content differentiates between different classes of instructions. Refer to Figure 25A-Figure 25B , the content of this field selects between Class A instructions and Class B instructions. In Figure 25A-Figure 25B , rounded rectangles are used to indicate that a particular value exists in the field (e.g., Class A 2568A and Class B 2568B for the class field 2568 in Figure 25A-Figure 25B respectively).

[0259] Class A instruction templates

[0260] In the case of the instruction template for Class A non-memory access 2505, the α field 2552 is interpreted as an RS field 2552A whose content differentiates which of different extended operation types is to be performed (e.g., the instruction templates for rounding-type operations 2510 without memory access and data transformation-type operations 2515 without memory access specify rounding 2552A.1 and data transformation 2552A.2, respectively), while the β field 2554 differentiates which of the operations of the specified type is to be performed. In the instruction template for non-memory access 2505, the scale field 2560, the displacement field 2562A, and the displacement factor field 2562B do not exist.

[0261] Instruction template for non-memory access - fully rounded control-type operation

[0262] In the instruction template for the fully rounded control-type operation 2510 without memory access, the β field 2554 is interpreted as a rounding control field 2554A whose (multiple) content provides static rounding. Although in the described embodiments of the present invention the rounding control field 2554A includes a suppress all floating-point exceptions (SAE) field 2556 and a rounding operation control field 2558, alternative embodiments may support both concepts, may encode both concepts into the same field, or may have only one or the other of these concepts / fields (e.g., may have only the rounding operation control field 2558).

[0263] SAE field 2556 - whose content differentiates whether to disable the reporting of exception events; when the content of the SAE field 2556 indicates enabling suppression, a given instruction does not report any kind of floating-point exception flag and does not invoke any floating-point exception handler.

[0264] Rounding operation control field 2558 - whose content differentiates which of a set of rounding operations is to be performed (e.g., round up, round down, round towards zero, and round to nearest). Thus, the rounding operation control field 2558 allows the rounding mode to be changed instruction by instruction. In one embodiment of the present invention in which the processor includes a control register for specifying the rounding mode, the content of the rounding operation control field 2550 overrides the register value.

[0265] Instruction template for non-memory access - data transformation-type operation

[0266] In the instruction template for the data transformation-type operation 2515 without memory access, the β field 2554 is interpreted as a data transformation field 2554B, whose content differentiates which of multiple data transformations is to be performed (e.g., no data transformation, mix, broadcast).

[0267] In the case of the instruction template for a class A memory access 2520, the α field 2552 is interpreted as an eviction hint field 2552B, the content of which differentiates which one of the eviction hints is to be used (in Figure 25A for the instruction template for memory access timeliness 2525 and the instruction template for memory access non - timeliness 2530, timeliness 2552B.1 and non - timeliness 2552B.2 are specified respectively), while the β field 2554 is interpreted as a data manipulation field 2554C, the content of which differentiates which one of multiple data manipulation operations (also called primitives) is to be performed (e.g., no manipulation, broadcast, up - conversion of the source, and down - conversion of the destination). The instruction template for the memory access 2520 includes a scale field 2560 and optionally includes a displacement field 2562A or a displacement factor field 2562B.

[0268] Vector memory instructions use conversion support to perform vector loads from memory and vector stores to memory. Like ordinary vector instructions, vector memory instructions transfer data to / from memory in a data - element - by - data - element manner, where the elements actually transferred are specified by the content of a vector mask selected as a write mask.

[0269] Instruction template for memory access - Timeliness

[0270] Timely data is data that may be reused quickly enough to benefit from the cache. However, this is a hint, and different processors can implement it in different ways, including completely ignoring the hint.

[0271] Instruction template for memory access - Non - timeliness

[0272] Non - timely data is data that cannot be reused quickly enough to benefit from the cache in the level 1 cache and should be given eviction priority. However, this is a hint, and different processors can implement it in different ways, including completely ignoring the hint.

[0273] Class B instruction template

[0274] In the case of the class B instruction template, the α field 2552 is interpreted as a write mask control (Z) field 2552C, the content of which differentiates whether the write mask operation controlled by the write mask field 2570 should be merge or zero.

[0275] In the case of the instruction template for a Class B non-memory access 2505, a portion of the β field 2554 is interpreted as an RL field 2557A, the content of which differentiates which one of different extended operation types is to be performed (e.g., for the instruction template of a partial rounding control type operation 2512 with write mask control for non-memory access and the instruction template of a VSIZE type operation 2517 with write mask control for non-memory access, rounding 2557A.1 and vector length (VSIZE) 2557A.2 are specified respectively), while the remaining portion of the β field 2554 differentiates which one of the operations of the specified type is to be performed. In the instruction template for a non-memory access 2505, the scale field 2560, the displacement field 2562A, and the displacement factor field 2562B do not exist.

[0276] In the instruction template for a partial rounding control type operation 2510 with write mask control for non-memory access, the remaining portion of the β field 2554 is interpreted as a rounding operation field 2559A, and the reporting of exception events is disabled (the given instruction does not report any kind of floating-point exception flag and does not invoke any floating-point exception handler).

[0277] The rounding operation control field 2559A - like the rounding operation control field 2558, the content of which differentiates which one of a set of rounding operations is to be performed (e.g., rounding up, rounding down, rounding towards zero, and rounding to nearest). Thus, the rounding operation control field 2559A allows the rounding mode to be changed instruction by instruction. In one embodiment of the present invention in which the processor includes a control register for specifying the rounding mode, the content of the rounding operation control field 2550 overrides the register value.

[0278] In the instruction template for a VSIZE type operation 2517 with write mask control for non-memory access, the remaining portion of the β field 2554 is interpreted as a vector length field 2559B, the content of which differentiates which one of multiple data vector lengths is to be performed (e.g., 128 bytes, 256 bytes, or 512 bytes).

[0279] In the case of the instruction template for a Class B memory access 2520, a portion of the β field 2554 is interpreted as a broadcast field 2557B, the content of which differentiates whether a broadcast type data manipulation operation is to be performed, while the remaining portion of the β field 2554 is interpreted as a vector length field 2559B. The instruction template for a memory access 2520 includes a scale field 2560 and optionally includes a displacement field 2562A or a displacement factor field 2562B.

[0280] For the general vector friendly instruction format 2500, the full opcode field 2574 is shown to include a format field 2540, a base operation field 2542, and a data element width field 2564. Although one embodiment is shown in which the full opcode field 2574 includes all of these fields, in embodiments that do not support all of these fields, the full opcode field 2574 includes less than all of these fields. The full opcode field 2574 provides the operation code (opcode).

[0281] The extended operation field 2550, the data element width field 2564, and the write mask field 2570 allow these features to be specified in the general vector friendly instruction format on a per-instruction basis.

[0282] The combination of the write mask field and the data element width field creates various types of instructions because these instructions allow the mask to be applied based on different data element widths.

[0283] The various instruction templates that occur within classes A and B are beneficial in different scenarios. In some embodiments of the present invention, different processors or different cores within a processor may support only class A, only support class B, or may support both classes. For example, a high-performance general out-of-order core intended for general computing may support only class B, a core intended primarily for graphics and / or scientific (throughput) computing may support only class A, and a core intended for both general computing and graphics and / or scientific (throughput) computing may support both class A and class B (of course, a core with some mix of templates and instructions from both classes, but not all templates and instructions from both classes, is within the scope of the present invention). Similarly, a single processor may include multiple cores, all of which support the same class, or where different cores support different classes. For example, in a processor with separate graphics and general cores, one core in the graphics core intended primarily for graphics and / or scientific computing may support only class A, while one or more in the general core may be high-performance general cores with out-of-order execution and register renaming intended for general computing that support only class B. Another processor without a separate graphics core may include one or more general in-order or out-of-order cores that support both class A and class B. Of course, in different embodiments of the present invention, features from one class may also be implemented in other classes. This will cause programs written in a high-level language to be (e.g., just-in-time compiled or statically compiled) into various different executable forms, which include: 1) a form having only the instructions of the (multiple) classes supported by the target processor for execution; or 2) a form having alternative routines and control flow code, where the alternative routines are written using different combinations of instructions from all classes, and the control flow code selects these routines for execution based on the instructions supported by the processor currently executing the code.

[0284] Exemplary Special Vector Friendly Instruction Format

[0285] Figure 26A is a block diagram illustrating an exemplary special vector friendly instruction format according to embodiments of the present invention. Figure 26A Illustrated is a special vector friendly instruction format 2600, which specifies positions, sizes, interpretations, and the order of fields, as well as the values of some of those fields, and in this sense the special vector friendly instruction format 800 is special. The special vector friendly instruction format 2600 can be used to extend the x86 instruction set, and thus some of the fields in this format are similar to or the same as those used in the existing x86 instruction set and its extensions (e.g., AVX). This format remains consistent with the prefix encoding field, real opcode byte field, MOD R / M field, SIB field, displacement field, and immediate number field of the existing x86 instruction set with extensions. Illustrated are fields from Figure 25A-Figure 25B and fields from Figure 26A are mapped to fields from Figure 25A-Figure 25B

[0286] It should be understood that although embodiments of the present invention are described with reference to the special vector friendly instruction format 2600 in the context of the general vector friendly instruction format 2500 for illustrative purposes, the present invention is not limited to the special vector friendly instruction format 2600, except where stated. For example, the general vector friendly instruction format 2500 contemplates various possible sizes for various fields, while the special vector friendly instruction format 2600 is shown with fields of specific sizes. As a specific example, although the data element width field 2564 is shown as a one-bit field in the special vector friendly instruction format 2600, the present invention is not limited thereto (i.e., the general vector friendly instruction format 2500 contemplates other sizes for the data element width field 2564).

[0287] The general vector friendly instruction format 2500 includes fields listed in the order shown below in Figure 26A

[0288] EVEX Prefix (bytes 0 - 3) 2602 - Encoded in four-byte form.

[0289] Format Field 2540 (EVEX byte 0, bits [7:0]) - The first byte (EVEX byte 0) is the format field 2540, and it contains 0x62 (a unique value used to distinguish the vector friendly instruction format in one embodiment of the present invention).

[0290] The second - fourth bytes (EVEX bytes 1 - 3) include multiple bit fields that provide special capabilities.

[0291] ​​REX field 2605 (EVEX byte 1, bits [7-5]) - consists of the EVEX.R bit field (EVEX byte 1, bit [7] – R), the EVEX.X bit field (EVEX byte 1, bit [6] – X), and (2557BEX byte 1, bit [5] – B). The EVEX.R, EVEX.X, and EVEX.B bit fields provide the same functionality as the corresponding VEX bit fields and are encoded in one's complement form, i.e., ZMM0 is encoded as 1111B and ZMM15 is encoded as 0000B. Other field pairs of these instructions encode the lower three bits (rrr, xxx, and bbb) of the register index known in the art such that Rrrr, Xxxx, and Bbbb can be formed by adding EVEX.R, EVEX.X, and EVEX.B.

[0292] REX' field 2510 - This is the first part of the REX' field 2510 and is the EVEX.R' bit field (EVEX byte 1, bit [4] – R') used to encode the higher 16 or lower 16 registers of the extended set of 32 registers. In one embodiment of the present invention, this bit is stored in a bit-reversed format together with other bits indicated below to distinguish (in the well-known x86 32-bit mode) from the BOUND instruction whose real opcode byte is 62 but does not accept the value 11 in the MOD field in the MODR / M field (described below); alternative embodiments of the present invention do not store this bit and other bits indicated below in a reversed format. The value 1 is used to encode the lower 16 registers. In other words, R'Rrrr is formed by combining EVEX.R', EVEX.R, and other RRRs from other fields.

[0293] Opcode mapping field 2615 (EVEX byte 1, bits [3:0] – mmmm) - its content encodes the implicit leading opcode byte (0F, 0F 38, or 0F 3).

[0294] Data element width field 2564 (EVEX byte 2, bit [7] – W) - denoted by the notation EVEX.W. EVEX.W is used to define the granularity (size) of the data type (32-bit data element or 64-bit data element).

[0295] EVEX.vvvv 2620 (EVEX byte 2, bits [6:3] - vvvv) - The functions of EVEX.vvvv may include the following: 1) EVEX.vvvv encodes the first source register operand specified in inverted (1's complement) form and is valid for instructions with two or more source operands; 2) EVEX.vvvv encodes the destination register operand specified in 1's complement form for a specific vector displacement; or 3) EVEX.vvvv does not encode any operand, this field is reserved, and shall contain 1111b. Thus, the EVEX.vvvv field 2620 encodes the 4 low-order bits of the first source register specifier stored in inverted (1's complement) form. Depending on the instruction, additional different EVEX bit fields are used to extend the specifier size to 32 registers.

[0296] EVEX.U 2568 class field (EVEX byte 2, bit [2] - U) - If EVEX.U = 0, it indicates class A or EVEX.U0; if EVEX.U = 1, it indicates class B or EVEX.U1.

[0297] Prefix encoding field 2625 (EVEX byte 2, bits [1:0] - pp) - Provides additional bits for the base operation field. In addition to supporting traditional SSE instructions in EVEX prefix format, this also has the benefit of compressing the SIMD prefix (the EVEX prefix only requires 2 bits instead of a byte to express the SIMD prefix). In one embodiment, to support traditional SSE instructions using SIMD prefixes (66H, F2H, F3H) in both traditional format and EVEX prefix format, these traditional SIMD prefixes are encoded into the SIMD prefix encoding field; and are expanded into traditional SIMD prefixes before being provided to the PLA at runtime (thus, without modification, the PLA can execute these traditional instructions in both traditional format and EVEX format). Although newer instructions may directly use the content of the EVEX prefix encoding field as an opcode extension, for consistency, a particular embodiment expands it in a similar manner but allows different meanings specified by these traditional SIMD prefixes. Alternative embodiments may redesign the PLA to support 2-bit SIMD prefix encoding and thus do not require expansion.

[0298] α field 2552 (EVEX byte 3, bit [7] – EH, also known as EVEX.EH, EVEX.rs, EVEX.RL, EVEX.write mask control, and EVEX.N; also shown as α) - As previously described, this field is context-specific.

[0299] β field 2554 (EVEX byte 3, bits [6:4] - SSS, also known as EVEX.s 2-0 、EVEX.r 2-0 、EVEX.rr1, EVEX.LL0, EVEX.LLB; also shown as βββ) - as previously described, this field is context - specific.

[0300] REX’ field 2510 - This is the remainder of the REX’ field and is the EVEX.V’ bit field (EVEX byte 3, bit [3] – V’) that can be used to encode the upper 16 or lower 16 registers of the extended 32 - register set. This bit is stored in bit - reversed format. The value 1 is used to encode the lower 16 registers. In other words, V’VVVV is formed by combining EVEX.V’ and EVEX.vvvv.

[0301] Write - mask field 2570 (EVEX byte 3, bits [2:0] - kkk) - Its content specifies the register index in the write - mask register, as previously described. In one embodiment of the present invention, the specific value EVEX.kkk = 000 has a special behavior that implies no write - mask is used for a particular instruction (this can be implemented in various ways, including using a write - mask hard - wired to all objects or hardware that bypasses the masking hardware).

[0302] Real opcode field 2630 (byte 4) is also called the opcode byte. A part of the opcode is specified in this field.

[0303] MOD R / M field 2640 (byte 5) includes the MOD field 2642, the Reg field 2644, and the R / M field 2646. As previously described, the content of the MOD field 2642 distinguishes between memory - access operations and non - memory - access operations. The role of the Reg field 2644 can be boiled down to two cases: encoding a destination register operand or a source register operand; or being regarded as an opcode extension and not being used to encode any instruction operand. The role of the R / M field 2646 can include the following: encoding an instruction operand that references a memory address; or encoding a destination register operand or a source register operand.

[0304] Scale, Index, Base (SIB) byte (byte 6) - As previously described, the content of the scale field 2550 is used for memory - address generation. SIB.xxx 2654 and SIB.bbb 2656 - The content of these fields has been previously mentioned for register indices Xxxx and Bbbb.

[0305] Displacement field 2562A (bytes 7 - 10) - When the MOD field 2642 contains 10, bytes 7 - 10 are the displacement field 2562A, and it works the same as a traditional 32 - bit displacement (disp32) and works at the byte granularity.

[0306] Displacement factor field 2562B (byte 7) - When the MOD field 2642 contains 01, byte 7 is the displacement factor field 2562B. The position of this field is the same as that of the traditional x86 instruction set 8 - bit displacement (disp8) which works at the byte granularity. Since disp8 is sign - extended, it can only address between - 128 and 127 byte offsets; in terms of a 64 - byte cache line, disp8 uses 8 bits that can be set to only four really useful values: - 128, - 64, 0, and 64; since a larger range is often needed, disp32 is used; however, disp32 requires 4 bytes. In contrast to disp8 and disp32, the displacement factor field 2562B is a reinterpretation of disp8; when using the displacement factor field 2562B, the actual displacement is determined by multiplying the content of the displacement factor field by the size (N) of the memory operand access. This type of displacement is called disp8*N. This reduces the average instruction length (one byte for the displacement but with a much larger range). Such compressed displacements are based on the assumption that the effective displacement is a multiple of the memory access granularity, and thus the redundant low - order bits of the address offset do not need to be encoded. In other words, the displacement factor field 2562B replaces the traditional x86 instruction set 8 - bit displacement. Thus, the displacement factor field 2562B is encoded in the same way as the x86 instruction set 8 - bit displacement (so, there is no change in the ModRM / SIB encoding rules), the only difference being that disp8 is overloaded to disp8*N. In other words, there is no change in the encoding rules or encoding length, but only a change in the hardware's interpretation of the displacement value (which requires scaling the displacement by the size of the memory operand to obtain a byte - based address offset). The immediate field 2572 operates as previously described.

[0307] Full opcode field

[0308] Figure 26B is a block diagram showing the fields that make up the full opcode field 2574 with a dedicated vector - friendly instruction format 2600 according to an embodiment of the present invention. Specifically, the full opcode field 2574 includes a format field 2540, a base operation field 2542, and a data element width (W) field 2564. The base operation field 2542 includes a prefix encoding field 2625, an opcode mapping field 2615, and a real opcode field 2630.

[0309] Register index field

[0310] Figure 26C is a block diagram showing the fields of a dedicated vector-friendly instruction format 2600 that make up the register index field 2544 according to an embodiment of the present invention. Specifically, the register index field 2544 includes a REX field 2605, a REX' field 2610, a MODR / M.reg field 2644, a MODR / M.r / m field 2646, a VVVV field 2620, an xxx field 2654, and a bbb field 2656.

[0311] Expansion operation field

[0312] Figure 26D is a block diagram showing the fields of a dedicated vector-friendly instruction format 2600 that make up the expansion operation field 2550 according to an embodiment of the present invention. When the class (U) field 2568 contains 0, it indicates EVEX.U0 (class A 2568A); when it contains 1, it indicates EVEX.U1 (class B 2568B). When U = 0 and the MOD field 2642 contains 11 (indicating no memory access operation), the α field 2552 (EVEX byte 3, bit [7] – EH) is interpreted as the rs field 2552A. When the rs field 2552A contains 1 (rounding 2552A.1), the β field 2554 (EVEX byte 3, bits [6:4] – SSS) is interpreted as the rounding control field 2554A. The rounding control field 2554A includes a one-bit SAE field 2556 and a two-bit rounding operation field 2558. When the rs field 2552A contains 0 (data transformation 2552A.2), the β field 2554 (EVEX byte 3, bits [6:4] – SSS) is interpreted as a three-bit data transformation field 2554B. When U = 0 and the MOD field 2642 contains 00, 01, or 10 (indicating a memory access operation), the α field 2552 (EVEX byte 3, bit [7] – EH) is interpreted as the eviction hint (EH) field 2552B and the β field 2554 (EVEX byte 3, bits [6:4] – SSS) is interpreted as a three-bit data manipulation field 2554C.

[0313] When U = 1, the α field 2552 (EVEX byte 3, bit [7] – EH) is interpreted as the write mask control (Z) field 2552C. When U = 1 and the MOD field 2642 contains 11 (indicating no memory access operation), a part of the β field 2554 (EVEX byte 3, bit [4] – S0) is interpreted as the RL field 2557A; when it contains 1 (rounding 2557A.1), the rest of the β field 2554 (EVEX byte 3, bits [6-5] – S 2-1) is interpreted as a rounding operation field 2559A, and when the RL field 2557A contains 0 (VSIZE2557.A2), the remainder of the β field 2554 (EVEX byte 3, bits [6-5] - S 2-1 ) is interpreted as a vector length field 2559B (EVEX byte 3, bits [6-5] – L 1-0 ). When U = 1 and the MOD field 2642 contains 00, 01, or 10 (indicating a memory access operation), the β field 2554 (EVEX byte 3, bits [6:4] – SSS) is interpreted as a vector length field 2559B (EVEX byte 3, bits [6-5] – L 1-0 ) and a broadcast field 2,557B (EVEX byte 3, bit [4] – B).

[0314] Exemplary Register Architecture

[0315] Figure 27 is a block diagram of a register architecture 2700 according to an embodiment of the present invention. In the illustrated embodiment, there are 32 vector registers 2710 that are 512 bits wide; these registers are referred to as zmm0 through zmm31. The lower 256 bits of the lower 16 zmm registers overlay the registers ymm0-16. The lower 128 bits of the lower 16 zmm registers (the lower 128 bits of the ymm registers) overlay the registers xmm0-15. The specialized vector-friendly instruction format 2600 operates on these overlaid register stacks as illustrated in the following table.

[0316]

[0317] In other words, the vector length field 2559B selects between a maximum length and one or more other shorter lengths, where each such shorter length is half of the previous length; and instruction templates that do not have a vector length field 2559B operate on the maximum vector length. Additionally, in one embodiment, the class B instruction templates of the specialized vector-friendly instruction format 2600 operate on packed or scalar single / double-precision floating-point data and packed or scalar integer data. A scalar operation is an operation performed on the lowest-order data element position in a zmm / ymm / xmm register; depending on the embodiment, the higher-order data element positions either remain the same as before the instruction or are zeroed.

[0318] Write Mask Registers 2715 - In the illustrated embodiment, there are eight write mask registers (k0 through k7), each write mask register being 64 bits in size. In an alternative embodiment, the write mask registers 2715 are 16 bits in size. As previously described, in one embodiment of the present invention, vector mask register k0 cannot be used as a write mask; when the encoding that normally indicates k0 is used as a write mask, it selects the hard - wired write mask 0xFFFF, effectively disabling write masking for that instruction.

[0319] General - Purpose Registers 2725 - In the illustrated embodiment, there are sixteen 64 - bit general - purpose registers that are used with existing x86 addressing modes to address memory operands. These registers are referenced by the names RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 through R15.

[0320] Scalar Floating - Point Stack Register File (x87 Stack) 2745, overlaid with MMX Packed - Integer Flat Register File 2750 - In the illustrated embodiment, the x87 stack is an eight - element stack for performing scalar floating - point operations on 32 / 64 / 80 - bit floating - point data using the x87 instruction - set extension; and the MMX registers are used to perform operations on 64 - bit packed - integer data and to save operands for some operations performed between MMX and XMM registers.

[0321] Alternative embodiments of the present invention may use wider or narrower registers. Additionally, alternative embodiments of the present invention may use more, fewer, or different register files and registers.

[0322] Exemplary Core Architecture, Processor, and Computer Architecture

[0323] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, implementations of such cores can include: 1) general-purpose in-order cores intended for general computing; 2) high-performance general-purpose out-of-order cores intended for general computing; 3) specialized cores intended primarily for graphics and / or scientific (throughput) computing. Implementations of different processors can include: 1) a CPU that includes one or more general-purpose in-order cores intended for general computing and / or one or more general-purpose out-of-order cores intended for general computing; and 2) a coprocessor that includes one or more specialized cores intended primarily for graphics and / or scientific (throughput). Such different processors result in different computer system architectures, which can include: 1) a coprocessor on a chip separate from the CPU; 2) a coprocessor in the same package as the CPU but on a separate die; 3) a coprocessor on the same die as the CPU (in which case, such a coprocessor is sometimes referred to as specialized logic or as a specialized core, such as integrated graphics and / or scientific (throughput) logic); and 4) a system-on-a-chip that can include the described CPU (sometimes referred to as (one or more) application cores or (one or more) application processors), the coprocessor described above, and additional functionality on the same die. Exemplary core architectures are then described, followed by exemplary processor and computer architectures.

[0324] Exemplary Core Architectures

[0325] In-Order and Out-of-Order Core Block Diagrams

[0326] Figure 28A is a block diagram showing an exemplary in-order pipeline and an exemplary register-renamed out-of-order issue / execution pipeline in accordance with embodiments of the present invention. Figure 28B is a block diagram showing an exemplary embodiment of an in-order architecture core and an exemplary register-renamed out-of-order issue / execution architecture core to be included in a processor in accordance with multiple embodiments of the present invention. Figure 28A-Figure 28B The solid boxes in show the in-order pipeline and in-order core, while the optionally added dashed boxes show the register-renamed, out-of-order issue / execution pipeline and core. Given that the in-order aspects are a subset of the out-of-order aspects, the out-of-order aspects will be described.

[0327] In Figure 28A the processor pipeline 2800 includes a fetch stage 2802, a length decoding stage 2804, a decoding stage 2806, an allocation stage 2808, a renaming stage 2810, a scheduling (also known as dispatch or issue) stage 2812, a register read / memory read stage 2814, an execution stage 2816, a write-back / memory write stage 2818, an exception handling stage 2822, and a commit stage 2824.

[0328] Figure 28BA processor core 2890 is shown that includes a front-end unit 2830 coupled to an execution engine unit 2850, and both the execution engine unit and the front-end unit are coupled to a memory unit 2870. The core 2890 can be a reduced instruction set computing (RISC) core, a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, the core 2890 can be a specialized core, such as, for example, a network or communication core, a compression engine, a coprocessor core, a general-purpose computing graphics processing unit (GPGPU) core, a graphics core, and so on.

[0329] The front-end unit 2830 includes a branch prediction unit 2832 coupled to an instruction cache unit 2834, the instruction cache unit 2834 is coupled to an instruction translation lookaside buffer (TLB) 2836, the instruction translation lookaside buffer 2836 is coupled to an instruction fetch unit 2838, and the instruction fetch unit 2838 is coupled to a decode unit 2840. The decode unit 2840 (or decoder) can decode the instruction and generate, as output, one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals that are decoded from, otherwise reflect, or are derived from the original instruction. The decode unit 2840 can be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, look-up tables, hardware implementations, programmable logic arrays (PLAs), microcode read-only memories (ROMs), etc. In one embodiment, the core 2890 includes a microcode ROM or other medium for storing the microcode of certain macroinstructions (e.g., in the decode unit 2840 or otherwise within the front-end unit 2830). The decode unit 2840 is coupled to a rename / allocator unit 2852 in the execution engine unit 2850.

[0330] The execution engine unit 2850 includes a rename / allocator unit 2852 that is coupled to a retirement unit 2854 and a collection 2856 of one or more scheduler units. The (multiple) scheduler units 2856 represent any number of different schedulers, including reservation stations, a central instruction window, and the like. The (multiple) scheduler units 2856 are coupled to the (multiple) physical register file units 2858. Each physical register file unit 2858 represents one or more physical register files, where different physical register files store one or more different data types, such as scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points, status (e.g., an instruction pointer that is the address of the next instruction to be executed), and the like. In one embodiment, the physical register file unit 2858 includes a vector register unit, a write mask register unit, and a scalar register unit. These register units may provide architectural vector registers, vector mask registers, and general-purpose registers. The (multiple) physical register file units 2858 are overlapped by the retirement unit 2854 to illustrate various ways in which register renaming and out-of-order execution can be implemented (e.g., using the (multiple) reorder buffers and the (multiple) retirement register files; using the (multiple) future files, the (multiple) history buffers, and the (multiple) retirement register files; using register maps and register pools; and so on). The retirement unit 2854 and the (multiple) physical register file units 2858 are coupled to the (multiple) execution clusters 2860. The (multiple) execution clusters 2860 include a collection of one or more execution units 2862 and a collection of one or more memory access units 2864. The execution units 2862 can perform various operations (e.g., shift, add, subtract, multiply) and can operate on various data types (e.g., scalar floating point, packed integer, packed floating point, vector integer, vector floating point). Although some embodiments may include several execution units dedicated to a particular function or set of functions, other embodiments may include only one execution unit or multiple execution units that all perform all functions. The (multiple) scheduler units 2856, the (multiple) physical register file units 2858, and the (multiple) execution clusters 2860 are shown as potentially multiple because certain embodiments create separate pipelines for certain types of data / operations (e.g., a scalar integer pipeline, a scalar floating point / packed integer / packed floating point / vector integer / vector floating point pipeline, and / or a memory access pipeline that each has its own scheduler unit, physical register file unit, and / or execution cluster - and in the case of a separate memory access pipeline, certain embodiments are implemented where only the execution cluster of that pipeline has the (multiple) memory access units 2864). It should also be understood that in the case of using separate pipelines, one or more of these pipelines can be out-of-order issue / execution, and the remaining pipelines can be in-order issue / execution.

[0331] A collection of memory access units 2864 is coupled to a memory unit 2870, which includes a data TLB unit 2872, which is coupled to a data cache unit 2874, which is coupled to a level 2 (L2) cache unit 2876. In one exemplary embodiment, the memory access units 2864 may include a load unit, a store address unit, and a store data unit, each of which is coupled to the data TLB unit 2872 in the memory unit 2870. An instruction cache unit 2834 is also coupled to the level 2 (L2) cache unit 2876 in the memory unit 2870. The L2 cache unit 2876 is coupled to one or more other levels of cache and ultimately to main memory.

[0332] As an example, an exemplary register-renaming out-of-order issue / execution core architecture may implement a pipeline 2800 as follows: 1) Instruction fetch 2838 performs a fetch stage 2802 and a length decoding stage 2804; 2) A decode unit 2840 performs a decode stage 2806; 3) A rename / allocator unit 2852 performs an allocation stage 2808 and a rename stage 2810; 4) A (multiple) scheduler unit 2856 performs a schedule stage 2812; 5) A (multiple) physical register file unit 2858 and a memory unit 2870 perform a register read / memory read stage 2814; An execution cluster 2860 performs an execution stage 2816; 6) The memory unit 2870 and the (multiple) physical register file units 2858 perform a write-back / memory write stage 2818; 7) Each unit may be involved in an exception handling stage 2822; and 8) A retirement unit 2854 and the (multiple) physical register file units 2858 perform a commit stage 2824.

[0333] The core 2890 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions added with more recent versions); the MIPS instruction set of MIPS Technologies, Inc., in Sunnyvale, California; the ARM instruction set of ARM Holdings plc, in Sunnyvale, California (with optional additional extensions such as NEON)), including the (multiple) instructions described herein. In one embodiment, the core 2890 includes logic for supporting SIMD instruction set extensions (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using SIMD data.

[0334] It should be understood that the core supports multithreading (a collection of two or more parallel operations or threads), and this multithreading can be accomplished in various ways, including time division multithreading, simultaneous multithreading (where a single physical core provides a logical core for each of the threads that the physical core is simultaneously multithreading), or a combination thereof (e.g., time division fetching and decoding and subsequent simultaneous multithreading such as in simultaneous multithreading in hyperthreading technology).

[0335] Although register renaming is described in the context of out-of-order execution, it should be understood that register renaming can be used in an in-order architecture. Although the illustrated embodiments of the processor also include separate instruction and data cache units 2834 / 2874 and a shared L2 cache unit 2876, alternative embodiments may have a single internal cache for both instructions and data, such as, for example, a level 1 (L1) internal cache or multiple levels of internal caches. In some embodiments, the system may include a combination of an internal cache and an external cache outside the core and / or processor. Alternatively, all caches may be outside the core and / or processor.

[0336] Specific exemplary in-order core architecture

[0337] Figure 29A-Figure 29B A block diagram illustrating a more specific exemplary in-order core architecture, which would be one of several logic blocks in a chip (including other cores of the same type and / or different types). Depending on the application, the logic blocks communicate with some fixed function logic, memory I / O interfaces, and other necessary I / O logic via a high bandwidth interconnect network (e.g., a ring network).

[0338] Figure 29A is a block diagram of a single processor core according to an embodiment of the present invention, the connection of the single processor core to the on-die interconnect network 2902, and a local subset 2904 of its level 2 (L2) cache. In one embodiment, the instruction decoder 2900 supports the x86 instruction set with a compact data instruction set extension. The L1 cache 2906 allows low latency access to cache memory for entry into scalar and vector units. Although in one embodiment (for simplicity of design), the scalar unit 2908 and the vector unit 2910 use separate register sets (scalar registers 2912 and vector registers 2914 respectively), and the data transferred between these registers is written to memory and then read back from the level 1 (L1) cache 2906, alternative embodiments of the present invention may use different methods (e.g., using a single register set or including a communication path that allows data to be transferred between the two register banks without being written and read back).

[0339] The local subset 2904 of the L2 cache is part of the global L2 cache, which is partitioned into multiple separate local subsets, one for each processor core. Each processor core has a direct access path to its own local subset 2904 of the L2 cache. Data read by a processor core is stored in its L2 cache subset 2904 and can be quickly accessed in parallel with other processor cores accessing their own local L2 cache subsets. Data written by a processor core is stored in the L2 cache subset 2904 of that processor core and is flushed from other subsets if necessary. The ring network ensures data consistency for shared data. The ring network is bi-directional to allow agents such as processor cores, L2 caches, and other logic blocks to communicate with each other within the chip. Each ring data path is 1012 bits wide for each direction.

[0340] Figure 29B is a partial expanded view of a processor core according to various embodiments of the present invention Figure 29A in Figure 29B The L1 data cache 2906A portion includes the L1 cache 2904, as well as more details regarding the vector unit 2910 and the vector registers 2914. Specifically, the vector unit 2910 is a 16-wide vector processing unit (VPU) (see the 16-wide ALU 2928) that executes one or more of integer, single-precision floating-point, and double-precision floating-point instructions. The VPU supports mixing of register inputs using the mixing unit 2920, numerical conversion using the numerical conversion units 2922A-B, and copying of memory inputs using the copy unit 2924. The write mask register 2926 allows assertion of the resulting vector writes.

[0341] Figure 30 is a block diagram of a processor 3000 that according to various embodiments of the present invention may have more than one core, may have an integrated memory controller, and may have an integrated graphics device. Figure 30 The solid box diagram of shows the processor 3000, which has a single core 3002A, a system agent 3010, and a set 3016 of one or more bus controller units, while the optional additional dashed box shows an alternative processor 3000, which has multiple cores 3002A-N, a set 3014 of one or more integrated memory controller units in the system agent unit 3010, and dedicated logic 3008.

[0342] Thus, different implementations of the processor 3000 can include: 1) a CPU, where the dedicated logic 3008 is integrated graphics and / or scientific (throughput) logic (which can include one or more cores), and the cores 3002A-N are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, a combination of both); 2) a coprocessor, where the cores 3002A-N are a large number of dedicated cores designed primarily for graphics and / or scientific (throughput); and 3) a coprocessor, where the cores 3002A-N are a large number of general-purpose in-order cores. Thus, the processor 3000 can be a general-purpose processor, a coprocessor, or a special-purpose processor, such as, for example, a network or communication processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput integrated many-core (MIC) coprocessor (including 30 or more cores), an embedded processor, and so on. The processor can be implemented on one or more chips. The processor 3000 can be a part of one or more substrates, and / or can be implemented on one or more substrates using any of a variety of process technologies, such as, for example, BiCMOS, CMOS, or NMOS.

[0343] The memory hierarchy includes one or more cache levels within the cores, a set of or one or more shared cache units 3006, and external memory (not shown) coupled to the set of integrated memory controller units 3014. The set of shared cache units 3006 can include one or more intermediate-level caches, a last-level cache (LLC), and / or a combination of the above, intermediate-level caches such as a level 2 (L2), level 3 (L3), level 4 (L4), or other level cache. Although in one embodiment, the ring-based interconnect unit 3012 interconnects the integrated graphics logic 3008 (the integrated graphics logic 3008 is an example thereof and is also referred to herein as dedicated logic), the set of shared cache units 3006, and the system agent unit 3010 / (multiple) integrated memory controller units 3014, alternative embodiments can use any number of well-known techniques to interconnect such units. In one embodiment, coherence is maintained between one or more cache units 3006 and the cores 3002-A-N.

[0344] In some embodiments, one or more of the cores 3002A-N are capable of implementing multithreading. The system agent 3010 includes those components that coordinate and operate the cores 3002A-N. The system agent unit 3010 can include, for example, a power control unit (PCU) and a display unit. The PCU can be the logic and components required to regulate the power states of the cores 3002A-N and the integrated graphics logic 3008, or can include such logic and components. The display unit is used to drive one or more externally connected displays.

[0345] The cores 3002A-N can be homogeneous or heterogeneous in terms of the architecture instruction set; that is, two or more of the cores 3002A-N may be capable of executing the same instruction set, while other cores may be capable of executing only a subset of the instruction set or a different instruction set.

[0346] Exemplary computer architecture

[0347] Figure 31-Figure 34 is a block diagram of an exemplary computer architecture. Other system designs and configurations known in the art for laptop devices, desktop computers, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular telephones, portable media players, handheld devices, and various other electronic devices are also suitable. In general, a wide variety of systems or electronic devices capable of incorporating a processor and / or other execution logic as disclosed herein are generally suitable.

[0348] Now refer to Figure 31 , shown is a block diagram of a system 3100 according to an embodiment of the present invention. The system 3100 may include one or more processors 3110, 3115, which are coupled to a controller hub 3120. In one embodiment, the controller hub 3120 includes a Graphics Memory Controller Hub (GMCH) 3190 and an Input / Output Hub (IOH) 3150 (which may be on separate chips); the GMCH 3190 includes a memory and graphics controller, to which a memory 3140 and a coprocessor 3145 are coupled; the IOH 3150 couples input / output (I / O) devices 3160 to the GMCH 3190. Alternatively, one or both of the memory and graphics controller are integrated within the processor (as described herein), and the memory 3140 and the coprocessor 3145 are directly coupled to the processor 3110 and the controller hub 3120, which is in a single chip with the IOH 3150.

[0349] The optionality of the additional processor 3115 is represented by a dashed line in Figure 31 . Each processor 3110, 3115 may include one or more of the processing cores described herein and may be a certain version of the processor 3000.

[0350] The memory 3140 can be, for example, a dynamic random access memory (DRAM), a phase change memory (PCM), or a combination of the two. For at least one embodiment, the controller hub 3120 communicates with the processors 3110, 3115 via a multi-branch bus such as a front side bus (FSB), a point-to-point interface such as a QuickPath Interconnect (QPI), or a similar connection 3195.

[0351] In one embodiment, the coprocessor 3145 is a specialized processor, such as, for example, a high throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, and the like. In one embodiment, the controller hub 3120 can include an integrated graphics accelerator.

[0352] There are various differences between the physical resources 3110, 3115 according to a metric spectrum including advantages such as architecture, microarchitecture, thermal, power consumption characteristics, and the like.

[0353] In one embodiment, the processor 3110 executes instructions that control general types of data processing operations. Coprocessor instructions can be embedded within these instructions. The processor 3110 identifies these coprocessor instructions as being of a type that should be executed by the attached coprocessor 3145. Accordingly, the processor 3110 issues these coprocessor instructions (or control signals representing coprocessor instructions) to the coprocessor 3145 on a coprocessor bus or other interconnect. The coprocessor(s) 3145 receive and execute the received coprocessor instructions.

[0354] Now referring Figure 32 , shown is a block diagram of a more specific first exemplary system 3200 in accordance with an embodiment of the present invention. As Figure 32 shown, the multi-processor system 3200 is a point-to-point interconnect system and includes a first processor 3270 and a second processor 3280 coupled via a point-to-point interconnect 3250. Each of the processors 3270 and 3280 can be a certain version of the processor 3000. In one embodiment of the present invention, the processors 3270 and 3280 are the processors 3110 and 3115 respectively, and the coprocessor 3238 is the coprocessor 3145. In another embodiment, the processors 3270 and 3280 are the processor 3110 and the coprocessor 3145 respectively.

[0355] Processors 3270 and 3280 are shown as each including an integrated memory controller (IMC) unit 3272 and 3282, respectively. Processor 3270 also includes point-to-point (P-P) interfaces 3276 and 3278 as part of its bus controller unit; similarly, second processor 3280 includes P-P interfaces 3286 and 3288. Processors 3270, 3280 may exchange information via P-P interface 3250 using point-to-point (P-P) interface circuits 3278, 3288. As Figure 32 shown, IMCs 3272 and 3282 couple the processors to respective memories, namely memories 3232 and 3234, which may be portions of main memories locally attached to the respective processors.

[0356] Processors 3270, 3280 may each exchange information with chipset 3290 via respective P-P interfaces 3252, 3254 using point-to-point interface circuits 3276, 3294, 3286, 3298. Chipset 3290 may optionally exchange information with coprocessor 3238 via high performance interface 3292. In one embodiment, coprocessor 3238 is a special-purpose processor such as, for example, a high throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, and the like.

[0357] A shared cache (not shown) may be included in either processor or external to both processors but connected to these processors via a P-P interconnect such that if a processor is placed in a low power mode, local cache information of either or both processors may be stored in the shared cache.

[0358] Chipset 3290 may be coupled to first bus 3216 via interface 3296. In one embodiment, first bus 3216 may be a Peripheral Component Interconnect (PCI) bus, or a bus such as a PCI Express bus or another third generation I / O interconnect bus, but the scope of the present invention is not limited thereto.

[0359] As Figure 32As shown, various I / O devices 3214 may be coupled to a first bus 3216 along with a bus bridge 3218 that couples the first bus 3216 to a second bus 3220. In one embodiment, one or more additional processors 3215, such as a coprocessor, a high-throughput MIC processor, a GPGPU, an accelerator (such as, for example, a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array, or any other processor, are coupled to the first bus 3216. In one embodiment, the second bus 3220 may be a low pin count (LPC) bus. Various devices may be coupled to the second bus 3220, which in one embodiment includes, for example, a keyboard / mouse 3222, a communication device 3227, and a storage unit 3228 such as a disk drive or other mass storage device that may include instructions / code and data 3230. Additionally, audio I / O 3224 may be coupled to the second bus 3220. Note that other architectures are possible. For example, instead of Figure 32 a point-to-point architecture, the system may implement a multi-branch bus or other such architecture.

[0360] Now referring to Figure 33 , shown is a block diagram of a more specific second exemplary system 3300 in accordance with an embodiment of the present invention. Figure 32 And Figure 33 Similar elements in Figure 33 are denoted with like reference numerals, and certain aspects of Figure 32 are omitted in Figure 33 to avoid obscuring other aspects of

[0361] Figure 33 It is shown that processors 3270, 3280 may each include integrated memory and I / O control logic ("CL") 3272 and 3282. Thus, CL 3272, 3282 includes an integrated memory controller unit and includes I / O control logic. Figure 33 It is illustrated that not only memories 3232, 3234 are coupled to CL 3272, 3282, but also I / O devices 3314 are coupled to control logic 3272, 3282. Conventional I / O devices 3315 are coupled to a chipset 3290.

[0362] Now referring to Figure 34 , shown is a block diagram of a SoC 3400 in accordance with an embodiment of the present invention. In Figure 30 , similar components have the same reference numerals. Additionally, the dashed boxes are optional features on a more advanced SoC. In Figure 34In [the figure], the (multiple) interconnect units 3402 are coupled to: an application processor 3410, which includes a set of one or more cores 3002A-N and (multiple) shared cache units 3006, and the set of one or more cores 3002A-N includes cache units 3004A-N; a system agent unit 3010; (multiple) bus controller units 3016; (multiple) integrated memory controller units 3014; a set of one or more coprocessors 3420, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random access memory (SRAM) unit 3430; a direct memory access (DMA) unit 3432; and a display unit 3440 for coupling to one or more external displays. In one embodiment, the (multiple) coprocessors 3420 include dedicated processors, such as, for example, a network or communication processor, a compression engine, a GPGPU, a high-throughput MIC processor, an embedded processor, and the like.

[0363] Embodiments of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of such implementations. Embodiments of the present invention may be implemented as a computer program or program code executed on a programmable system that includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0364] Program code (such as, Figure 32 the code 3230 illustrated in [the figure]) may be applied to input instructions to perform the functions described herein and generate output information. The output information may be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor, such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0365] The program code may be implemented in a high-level procedural programming language or an object-oriented programming language in order to communicate with the processing system. If desired, the program code may also be implemented in assembly language or machine language. In fact, the mechanisms described herein are not limited to any specific programming language scope. In any case, the language may be a compiled language or an interpreted language.

[0366] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium that represent various logic in a processor, which, when read by the machine, cause the machine to fabricate logic for performing the techniques described herein. Such representations, referred to as “IP cores,” may be stored on a tangible machine-readable medium and supplied to various customers or production facilities to be loaded into a manufacturing machine that actually fabricates the logic or processor.

[0367] Such machine-readable storage media may include, but are not limited to, non-transitory tangible arrangements of articles manufactured or formed by a machine or device, which include storage media such as: hard disks; any other type of disk, including floppy disks, optical disks, compact disk read-only memory (CD-ROM), compact disk rewritable (CD-RW), and magneto-optical disks; semiconductor devices such as read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM) and static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM); phase change memory (PCM); magnetic or optical cards; or any other type of medium suitable for storing electronic instructions.

[0368] Accordingly, embodiments of the present invention also include non-transitory tangible machine-readable media that contain instructions or contain design data, such as a hardware description language (HDL), that define the structures, circuits, devices, processors, and / or system features described herein. Such embodiments may also be referred to as program products.

[0369] Emulation (including binary translation, code morphing, etc.)

[0370] In some cases, an instruction converter may be used to convert instructions from a source instruction set to a target instruction set. For example, the instruction converter may transform (e.g., using static binary translation, dynamic binary translation including dynamic compilation), morph, emulate, or otherwise convert an instruction or instructions into one or more other instructions to be processed by a core. The instruction converter may be implemented in software, hardware, firmware, or a combination thereof. The instruction converter may be on the processor, off the processor, or partly on the processor and partly off the processor.

[0371] Figure 35 is a block diagram of an example of using a software instruction converter to convert binary instructions in a source instruction set into binary instructions in a target instruction set in accordance with an embodiment of the present invention. In the illustrated embodiment, the instruction converter is a software instruction converter, but alternatively, the instruction converter may be implemented in software, firmware, hardware, or various combinations thereof. Figure 35It is shown that a program using a high-level language 3502 can be compiled using an x86 compiler 3504 to generate x86 binary code 3506 that can be natively executed by a processor 3516 having at least one x86 instruction set core. The processor 3516 having at least one x86 instruction set core represents any processor that can perform substantially the same functions as an Intel processor having at least one x86 instruction set core by compatibly executing or otherwise processing: 1) an essential part of the instruction set of the Intel x86 instruction set core, or 2) a target code version of an application or other software targeted to run on an Intel processor having at least one x86 instruction set core to achieve substantially the same results as an Intel processor having at least one x86 instruction set core. The x86 compiler 3504 represents a compiler operable to generate x86 binary code 3506 (e.g., target code) that can be executed on a processor 3516 having at least one x86 instruction set core with or without additional linking processing. Similarly, Figure 35 It is shown that a program using a high-level language 3502 can be compiled using an alternative instruction set compiler 3508 to generate alternative instruction set binary code 3510 that can be natively executed by a processor 3514 not having at least one x86 instruction set core (e.g., a processor having cores that execute the MIPS instruction set of MIPS Technologies, Inc. of Sunnyvale, California and / or the ARM instruction set of ARM Holdings plc of Sunnyvale, California). An instruction converter 3512 is used to convert the x86 binary code 3506 into code that can be natively executed by a processor 3514 not having an x86 instruction set core. The converted code is not likely to be the same as the alternative instruction set binary code 3510 because instruction converters capable of doing so are difficult to manufacture; however, the converted code will perform general operations and be composed of instructions from an alternative instruction set. Thus, the instruction converter 3512 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device not having an x86 instruction set processor or core to execute the x86 binary code 3506 through emulation, simulation, or any other process.

Claims

1. A processor, comprising: Slice storage for storing a first two-dimensional slice and a second two-dimensional slice; Fetch circuit for fetching instructions; Decode circuit coupled to the fetch circuit, the decode circuit for decoding the instruction, the instruction having an opcode and indicating the first two-dimensional slice and the second two-dimensional slice; and Execution circuit coupled to the decode circuit and the slice storage, the execution circuit for: If a parameter associated with the second two-dimensional slice has a first value, zero out each element of the second two-dimensional slice.

2. The processor according to claim 1, wherein, The first value is a true value.

3. The processor according to claim 1 or 2, wherein, The parameter is a single bit.

4. The processor according to claim 1, wherein, The parameter is one of a plurality of parameters each corresponding to a different slice among a plurality of slices.

5. The processor according to claim 1, wherein, The parameter is a single bit of a byte, and wherein each bit of the byte corresponds to a different slice among a plurality of slices.

6. The processor according to claim 1, wherein, The execution circuit for: if the parameter has the first value, zero out each element of the first two-dimensional slice.

7. The processor according to claim 1, wherein, The slice storage includes a plurality of registers.

8. The processor according to claim 1, wherein, The second two-dimensional slice has configurable number of rows and columns.

9. The processor according to claim 1, further comprising a register for storing a value indicating the number of columns of the second two-dimensional sheet.

10. The processor according to claim 1, wherein, The number of rows of the second two-dimensional slice is based on the size of the data elements of the second two-dimensional slice.

11. The processor according to claim 1, wherein, The execution circuit for zeroing out each element of the second two-dimensional slice for: zeroing out the elements of each row of the second two-dimensional slice one by one.

12. A processor, comprising: Slice storage for storing a first two-dimensional slice and a second two-dimensional slice; Fetch circuit for fetching instructions; Decode circuit coupled to the fetch circuit, the decode circuit for decoding the instruction, the instruction having an opcode and indicating the first two-dimensional slice and the second two-dimensional slice; And Execution circuit coupled to the decode circuit and the slice storage, the execution circuit for: If a parameter has a first value, zero out each element of the first two-dimensional slice; and If the parameter has the first value, zero out each element of the second two-dimensional slice.

13. The processor according to claim 12, wherein, The parameter is one of a plurality of parameters at least including a second parameter corresponding to a third two-dimensional slice.

14. The processor according to claim 12 or 13, wherein, The parameter is a part of a byte.

15. The processor according to claim 12, wherein, The slice storage includes a plurality of registers.

16. The processor according to claim 12, wherein, The second two-dimensional slice has configurable number of rows and columns.

17. The processor according to claim 12, further comprising a register for storing a value indicating the number of columns of the second two-dimensional sheet.

18. The processor according to claim 12, wherein The number of rows of the second two-dimensional slice is based on the size of the data elements of the second two-dimensional slice.

19. The processor according to claim 12, wherein The execution circuit for zeroing out each element of the second two-dimensional slice for: zeroing out the elements of each row of the second two-dimensional slice one by one.

20. A method for a computer processor, comprising: Store a first two-dimensional slice and a second two-dimensional slice; Fetch an instruction; Decode the instruction, the instruction having an opcode and indicating the first two-dimensional slice and the second two-dimensional slice; and Execute the decoded instruction to: If a parameter associated with the second two-dimensional slice has a first value, zero out each element of the second two-dimensional slice.

21. A method for a computer processor, comprising: Store a first two-dimensional slice and a second two-dimensional slice; Fetch an instruction; Decode the instruction, the instruction having an opcode and indicating the first two-dimensional slice and the second two-dimensional slice; and Execute the decoded instruction to: If a parameter has a first value, zero out each element of the first two-dimensional slice; and If the parameter has the first value, zero out each element of the second two-dimensional slice.

22. A machine-readable medium comprising instructions that, when executed by one or more machines, cause the one or more machines to perform the method according to claim 20 or 21.

Citation Information

Patent Citations

  • Hardware for performing arithmetic operations

    CN102918495A

  • Optimizing register initialization operations

    CN103377037A