Systems, apparatus, and methods for fused multiply-add

By using the fused multiplication and accumulation instructions, using operands and accumulators of different sizes, the data size management problem in multiplication and accumulation operations is solved, and efficient and accurate calculation is achieved, avoiding overflow or saturation problems.

CN113885833BActive Publication Date: 2025-05-13INTEL CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111331383.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2016-10-20
Publication Date
2025-05-13
Estimated Expiration
2036-10-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively manage data size when dealing with multiplication and accumulation operations, resulting in the possible overflow or saturation in the calculation.

Method used

The fused multiplication accumulation instruction is adopted to achieve efficient accumulation of multiplication results by using operands and accumulators of different sizes, and saturation is performed when necessary.

Benefits of technology

The data size is effectively managed, overflow or saturation problems in calculations are avoided, and the accuracy and efficiency of multiplication and accumulation operations are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113885833B_ABST
    Figure CN113885833B_ABST
Patent Text Reader

Abstract

The present application discloses systems, apparatus, and methods for fused multiply-add. In some embodiments, the packed data elements of the first and second packed data source operands have a first size that is different from a second size of the packed data elements of the third packed data source operand. The execution circuit executes a decoded single instruction to perform, for each packed data element location of the destination operand: a multiplication of M packed data elements of N size from the first and second packed data sources corresponding to the packed data element location of the third packed data source, adding the results from the multiplications to the full-size packed data elements at the packed data element location of the third packed data source, and storing the addition results in the packed data element location destination corresponding to the packed data element location of the third packed data source, where M is equal to the full-size packed data element divided by N.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with PCT international application number PCT / US2016 / 057991, international application date October 20, 2016, application number 201680089435.5 entering the Chinese national phase, and entitled "System, device and method for fused multiplication and addition". Technical Field

[0002] The field of the invention relates generally to computer processor architecture, and more particularly to instructions that cause certain results when executed. Background Art

[0003] A common operation in linear algebra is the multiply-accumulate operation (e.g., c=c+a*b). Multiply-accumulate is typically a sub-operation in a stream of operations, such as a dot product between two vectors, or a single product of columns and rows in a matrix multiplication. For example,

[0004] C=0

[0005] For(I)

[0006] C+=A[l]*B[l]. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The present invention is illustrated by way of example and not limitation in the figures of the accompanying drawings in which like references indicate similar elements and in which:

[0008] Figure 1 illustrates an exemplary execution of a fused multiply-accumulate instruction using operands of different sizes according to an embodiment;

[0009] Figure 2 illustrates a power-of-two sized SIMD implementation in accordance with an embodiment, where the accumulator uses a larger input size than the input to the multiplier;

[0010] Figure 3 An embodiment of hardware for processing instructions such as fused multiply-accumulate instructions is illustrated;

[0011] Figure 4 An embodiment of a method performed by a processor to process a fused multiply-accumulate instruction is illustrated;

[0012] Figure 5 An embodiment of a subset of the execution of a fused multiply-accumulate is illustrated;

[0013] Figure 6 An embodiment of a pseudo code for implementing the instruction in hardware is illustrated;

[0014] Figure 7An embodiment of a subset of the execution of a fused multiply-accumulate is illustrated;

[0015] Figure 8 An embodiment of a pseudo code for implementing the instruction in hardware is illustrated;

[0016] Fig. 9 An embodiment of a subset of the execution of a fused multiply-accumulate is illustrated;

[0017] Fig.10 An embodiment of a pseudo code for implementing the instruction in hardware is illustrated;

[0018] Fig.11 An embodiment of a subset of the execution of a fused multiply-accumulate is illustrated;

[0019] Fig.12 An embodiment of a pseudo code for implementing the instruction in hardware is illustrated;

[0020] Fig.13A is a block diagram illustrating a general vector-friendly instruction format and a class A instruction template thereof according to an embodiment of the present invention;

[0021] Fig. 13B is a block diagram illustrating a general vector-friendly instruction format and a class B instruction template thereof according to an embodiment of the present invention;

[0022] Fig.14A is a block diagram illustrating an exemplary specific vector friendly instruction format according to an embodiment of the present invention;

[0023] Fig. 14B is a block diagram illustrating the fields of a particular vector friendly instruction format that make up a full opcode field in accordance with one embodiment of the present invention;

[0024] Fig. 14C is a block diagram illustrating fields of a particular vector friendly instruction format that constitute a register index field according to one embodiment of the present invention;

[0025] Fig.14D is a block diagram illustrating fields of a particular vector friendly instruction format that constitute an augment operation field according to one embodiment of the present invention;

[0026] Fig.15 is a block diagram of a register architecture according to one embodiment of the present invention;

[0027] Fig.16A is a block diagram illustrating both an exemplary in-order pipeline and an exemplary register renaming, out-of-order issue / execution pipeline according to an embodiment of the present invention;

[0028] Fig. 16B is a block diagram illustrating an exemplary embodiment of both an in-order architecture core and an exemplary register renaming, out-of-order issue / execution architecture core to be included in a processor in accordance with an embodiment of the present invention;

[0029] Fig.17A is a block diagram of a single processor core along with its connection to an on-die interconnect network 1702 and along with a local subset of its level 2 (L2) cache 1704 in accordance with an embodiment of the present invention;

[0030] Fig. 17B According to an embodiment of the present invention Fig.17A An expanded view of a portion of a processor core in FIG.

[0031] Fig.18 is a block diagram of a processor 1800 that may have more than one core, may have an integrated memory controller, and may have integrated graphics according to an embodiment of the present invention;

[0032] Fig.19 A block diagram of a system according to an embodiment of the present invention is shown;

[0033] Fig. 20 is a block diagram of a first more specific exemplary system according to an embodiment of the present invention;

[0034] Fig.21 is a block diagram of a second more specific exemplary system according to an embodiment of the present invention;

[0035] Fig. 22 is a block diagram of a SoC according to an embodiment of the present invention; and

[0036] Fig.23 is a block diagram comparing the use of a software instruction converter to convert binary instructions in a source instruction set to binary instructions in a target instruction set in accordance with an embodiment of the present invention. DETAILED DESCRIPTION

[0037] In the following description, numerous specific details are set forth. However, it is understood that embodiments of the present invention can be practiced without these specific details. In other instances, well-known circuits, structures, and techniques are not shown in detail in order not to obscure the understanding of this description.

[0038] References in this specification to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an embodiment, it is claimed that it is within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in conjunction with other embodiments, whether or not explicitly described.

[0039] In processing large data sets, memory and computation density can be increased by sizing the data types as small as possible. If the input terms come from sensor data, then 8-bit or 16-bit integer data can be expected as input. Neural network calculations, which can also be encoded to match this dense format, typically have "small" numbers as input terms. However, the accumulator sums the products, meaning that the accumulator should tolerate twice the number of bits in the input terms (the nature of multiplication) and potentially much more, in order to avoid overflow or saturation at any point in the calculation.

[0040] Detailed herein are embodiments that attempt to keep input data sizes small and aggregate to larger accumulators in a chain of fused multiply-accumulate (FMA) operations. Figure 1 An exemplary execution of a fused multiply-accumulate instruction using operands of different sizes according to an embodiment is illustrated. A first source 101 (e.g., a SIMD or vector register) and a second source 103 store "half-sized" packed data elements (e.g., a single-input, multiple-data (SIMD) or vector register) with respect to a third source 105 that stores full-sized packed data elements for accumulation. Any set of values ​​where the packed data element size is in this manner is supportable.

[0041] As shown, the values ​​stored in the same located packed data elements of the first and second sources 101 and 103 are multiplied together. For example, A0*B0, A1*B1, etc. The result of two such "half-sized" packed data element multiplications is added to the corresponding "full-sized" packed data element from the third source 105. For example, A0*B0+A1*B1+C0, etc.

[0042] The result is stored in a destination 107 (eg, a SIMD register) having a packed data element size of at least "full size." In some embodiments, the third source 105 and the destination 107 are identical.

[0043] Figure 2Illustrated is a power-of-two sized SIMD implementation in accordance with an embodiment, where the accumulator uses a larger input size than the input to the multiplier. Note that the source (for the multiplier) as well as the accumulator values ​​can be signed or unsigned values. For an accumulator with a 2X input size (in other words, the accumulator input value is twice the size of the packed data element size of the source), Table 201 illustrates different configurations. For byte sized sources, the accumulator uses word or half-precision floating point (HPFP) values, which are 16 bits in size. For word sized sources, the accumulator uses 32-bit integer or single-precision floating point (SPFP) values, which are 32 bits in size. For SPFP or 32-bit integer sized sources, the accumulator uses 64-bit integer or double-precision floating point (DPFP) values, which are 64 bits in size. Using Figure 1 As an example, when the packed data element size of Source 1 101 and Source 2 103 is 8 bits, then the accumulator will use 16-bit size data elements from Source 3 103. When the packed data element size of Source 1 101 and Source 2 103 is 16 bits, then the accumulator will use 32-bit size data elements from Source 3 103. When the packed data element size of Source 1 101 and Source 2 103 is 32 bits, then the accumulator will use 64-bit size data elements from Source 3 103.

[0044] Table 203 illustrates different configurations for an accumulator with a 4X input size (in other words, the accumulator input value is four times the size of the packed data element size of the source). For byte-sized sources, the accumulator uses 32-bit integers or single-precision floating point (SPFP) values, which are 32 bits in size. For word-sized sources, the accumulator uses 64-bit integers or double-precision floating point (DPFP) values, which are 64 bits in size. Figure 1 As an example, when the packed data element size of source 1 101 and source 2 103 is 8 bits, then the accumulator will use 32-bit sized data elements from source 3 103. When the packed data element size of source 1 101 and source 2 103 is 16 bits, then the accumulator will use 64-bit sized data elements from source 3 103.

[0045] Table 205 illustrates a configuration for an accumulator with an 8X input size (in other words, the accumulator input value is eight times the size of the packed data element size of the source). For byte-sized sources, the accumulator uses 64-bit integers or double-precision floating point (DPFP) values, which are 64 bits in size. Figure 1 As an example, when the packed data element size of source 1 101 and source 2 103 is 8 bits, then the accumulator will use a 64-bit size of data elements from source 3 103 .

[0046] Detailed herein are embodiments of instructions and circuits for fused multiply-accumulate. In some embodiments, the fused multiply-accumulate instruction has mixed precision and / or uses level reduction, as detailed herein.

[0047] Detailed herein are embodiments of instructions that, when executed, cause multiplication of M N-sized packed data elements from first and second sources corresponding to a packed data element location of a third source for each packed data element location of a destination, and add the results from the multiplications to the full-sized (relative to the N-sized packed data element) packed data element of the packed data element location of the third source, and store the results of the addition(s) in the packed data element location destination corresponding to the packed data element location of the third source, where M equals the full-sized packed data element divided by N. For example, when M equals 2 (e.g., the full-sized packed data element is 16 bits and N is 8 bits), successive packed data elements from the first source are multiplied to corresponding successive packed data elements of the second source.

[0048] Thus, detailed herein are embodiments of instructions that, when executed, cause multiplication of pairs of half-sized packed data elements from first and second sources, and add the results from these multiplications to full-sized (relative to half-sized packed data elements) packed data elements of a third source and store the results in a destination. In other words, in some embodiments, for each data element position i of the third source, there is a multiplication of data from data element position [2i] of the first source with data from data element position [2i] of the second source to generate a first result, a multiplication of data from data element position [2i+1] of the first source with data from data element position [2i+1] of the second source to generate a second result, and the first and second results are added to data from data element position i of the third source. In some embodiments, saturation is performed at the end of the addition. In some embodiments, the data from the first and / or second source is sign extended prior to multiplication.

[0049] Furthermore, detailed herein are embodiments of instructions that, when executed, cause multiplications of quartets of quarter-sized packed data elements from first and second sources, and add the results from these multiplications to full-sized (relative to the quarter-sized packed data elements) packed data elements of a third source and store the results in a destination. In other words, in some embodiments, for each data element position i of the third source, there is a multiplication of the data from the first source at data element position [4i] by the data from the second source at data element position [4i] to generate a first result, a multiplication of the data from the first source at data element position [4i+1] by the data from the second source at data element position [4i+1] to generate a second result, a multiplication of the data from the first source at data element position [4i+2] by the data from the second source at data element position [4i+2] to generate a second result, a multiplication of the data from the first source at data element position [4i+3] by the data from the second source at data element position [4i+3] to generate a second result, and the first, second, third, and fourth results are added to the data from the third source at data element position i. In some embodiments, saturation is performed at the end of the addition. In some embodiments, the data from the first and / or second source is sign extended prior to the multiplication.

[0050] In some embodiments of the integer version of the instruction, a saturation circuit is used to keep the sign of the operand when addition results in a too large value. In particular, saturation evaluation occurs on the infinite precision result between multi-way addition and writing to the destination. There are instances in which the largest positive number or the smallest negative number cannot be trusted because it may reflect that the calculation exceeds the container space. However, this can at least be checked. When the accumulator is a floating point and the input items are integers, the question to be answered is how and when to perform the conversion from the integer product so that there is no double rounding (double-rounding) from the partial item to the final floating point accumulation. In some embodiments, the product sum and the floating point accumulator are converted into infinite precision values ​​(hundreds of fixed-point numbers), addition is performed, and then a single rounding to the actual accumulator type is performed.

[0051] In some embodiments, when the input items are floating point operands, rounding and dealing with special values ​​(infinity and not a number (NAN)), the ordering of faults in the calculation needs to be addressed in the definition. In some embodiments, the order of operations is specified, which is simulated and ensures that the implementation delivers faults in this order. It may be impossible to avoid multiple roundings in the calculation process for such an implementation. Single-precision multiplication can be completely filled into the double-precision result, regardless of the input value. However, the horizontal addition of two such operations may not be suitable for double precision without rounding, and the sum may not be suitable for the accumulator without additional rounding. In some embodiments, rounding is performed during horizontal summation and performed once during accumulation.

[0052] Figure 3 An embodiment of hardware for processing instructions such as fused multiply-accumulate instructions is illustrated. As illustrated, a storage device 303 stores a fused multiply-accumulate instruction 301, which, to be executed, causes, for each packed data element location of a destination, a multiplication of M N-sized packed data elements from first and second sources corresponding to a packed data element location of a third source, adding the results from these multiplications to the full-sized (relative to the N-sized packed data elements) packed data elements of the packed data element location of the third source, and storing the results of the addition(s) in the packed data element location destination corresponding to the packed data element location of the third source, where M equals the full-sized packed data element divided by N.

[0053] An instruction 301 is received by a decode circuit 305. For example, the decode circuit 305 receives the instruction from a fetch logic / circuitry. The instruction includes fields for a first, second, and third source and a destination. In some embodiments, the source and destination are registers. Additionally, in some embodiments, the third source and destination are the same. The opcode and / or prefix of the instruction 301 includes an indication of the source and destination data element sizes {B / W / D / Q} in bytes, words, double words, and quad words, and the number of iterations.

[0054] A more detailed embodiment of at least one instruction format will be described later. Decoding circuit 305 decodes the instruction into one or more operations. In some embodiments, the decoding includes generating a plurality of micro-operations to be executed by an execution circuit (such as execution circuit 311). Decoding circuit 305 also decodes instruction prefixes.

[0055] In some embodiments, register renaming, register allocation and / or scheduling circuitry 307 provides functionality for one or more of: 1) renaming logical operand values ​​to physical operand values ​​(e.g., a register aliasing table in some embodiments), 2) assigning status bits and flags to decoded instructions, and 3) scheduling decoded instructions for execution on execution circuitry outside of an instruction pool (e.g., by using a reservation station in some embodiments). Registers (register files) and / or memory 308 store data, such as operands for instructions to be operated on by execution circuitry 309. Exemplary register types include packed data registers, general purpose registers, and floating point registers.

[0056] Execution circuit 309 executes the decoded instructions.

[0057] In some embodiments, retirement / writeback circuitry 311 architecturally commits the destination register to a register or memory and retires the instruction.

[0058] An embodiment of the format for a fused multiply-accumulate instruction is FMA[SOURCESIZE{B / W / D / Q}][DESTSIZE{B / W / D / Q}]DSTREG, SRC1, SRC2, SRC3. In some embodiments, FMA[SOURCESIZE{B / W / D / Q}][DESTSIZE{B / W / D / Q}] is the opcode and / or prefix of the instruction. B / W / D / Q indicates the data element size of the source / destination as byte, word, doubleword, and quadword. DSTREG is a field for a packed data destination register operand. SRC1, SRC2, and SRC3 are fields for a source such as a packed data register and / or memory.

[0059] An embodiment of the format for a fused multiply-accumulate instruction is FMA [SOURCESIZE {B / W / D / Q}] [DESTSIZE {B / W / D / Q}] DSTREG / SRC3, SRC1, SRC2. In some embodiments, FMA [SOURCESIZE {B / W / D / Q}] [DESTSIZE {B / W / D / Q}] is the opcode and / or prefix of the instruction. B / W / D / Q indicates the data element size of the source / destination as byte, word, doubleword, and quadword. DSTREG / SRC3 is a field for a packed data destination register operand and a third source operand. SRC1, SRC2, and SRC3 are fields for a source such as a packed data register and / or memory.

[0060] In some embodiments, the fused multiply-accumulate instruction includes a field for a write mask register operand (k) (e.g., FMA [SOURCESIZE {B / W / D / Q}] [DESTSIZE {B / W / D / Q}] {k} DSTREG / SRC3, SRC1, SRC2 or FMA [SOURCESIZE {B / W / D / Q}] [DESTSIZE {B / W / D / Q}] {k} DSTREG, SRC1, SRC2, SRC3). The write mask is used to conditionally control the element-wise operation and the updating of the result. Depending on the implementation, the write mask uses merge or zero masking. Instructions encoded with a predicted (write mask, write mask or k register) operand use the operand to conditionally control the element-wise calculation operation and the updating of the result to the destination operand. The predicted operand is known as an operation mask (write mask) register. The operation mask is a set of architectural registers of size MAX_KL (64 bits). Note that from this set of architectural registers, only k1 to k7 can be addressed as a prediction operand. k0 can be used as a conventional source or destination, but cannot be encoded as a prediction operand. It is also noted that the prediction operand can be used to enable memory fault suppression for certain instructions with memory operands (source or destination). As a prediction operand, the operation mask register contains one bit to control the operation / update of each data element of the vector register. Typically, the operation mask register can support instructions with the following element sizes: single precision, floating point (float 32), integer double word (integer 32), double precision floating point (float 64), integer quad word (integer 64). The length MAX_XL of the operation mask register is sufficient to handle up to 64 elements with one bit per element, i.e. 64 bits. For a given vector length, each instruction only accesses the number of the least significant mask bits required based on its data type. The operation mask register affects instructions at the granularity of the element. Therefore, any numeric or non-numeric operation of each data element and the element-wise update of the intermediate results to the destination operand are predicted on the corresponding bits of the operation mask register. In most embodiments, the operation mask used as the prediction operand obeys the following properties: 1) If the corresponding operation mask bit is not set, the operation of the instruction is not performed for the element (this means that no exceptions or violations can be caused by the operation on the masked-off element, and therefore, no exception flags are updated as a result of the mask operation); 2) If the corresponding write mask bit is not set, the destination element is not updated with the result of the operation. Instead, the destination element value must be maintained (merge-masked) or it must be zeroed (zero-masked); 3) For some instructions with memory operands, memory faults are suppressed for elements with 0 mask bits.Note that this feature provides a general construct for implementing control flow prediction, since the mask actually provides merging behavior for vector register destinations. As an alternative, masking can be used for zeroing, rather than merging, so that masked elements are updated with 0, rather than keeping the old value. The zeroing behavior is provided to remove the implicit dependency on the old value when it is not needed.

[0061] In an embodiment, the encoding of the instruction includes a memory addressing operand of a scale-index-base (SIB) type, which indirectly identifies a plurality of indexed destination locations in the memory. In one embodiment, a memory operand of the SIB type may include an encoding that identifies a base address register. The content of the base address register may represent a base address in the memory, and the address of a specific destination location in the memory is calculated according to the base address. For example, the base address may be the address of the first position in a block of potential destination locations for an extended vector instruction. In one embodiment, a memory operand of the SIB type may include an encoding that identifies an index register. Each element of the index register may specify an index or offset value that may be used to calculate the address of a corresponding destination location within a block of potential destination locations according to the base address. In one embodiment, a memory operand of the SIB type may include an encoding that specifies a scaling factor to be applied to each index value when calculating the corresponding destination address. For example, if a scaling factor value of four is encoded in a memory operand of the SIB type, each index value obtained from an element of the index register may be multiplied by four and then added to the base address to calculate the destination address.

[0062] In one embodiment, a memory operand of a SIB type with the form vm32{x,y,z} can identify a vector array of memory operands specified by using a memory addressing of the SIB type. In this example, the array of memory addresses is specified by using the following: a common base register, a constant scaling factor, and a vector index register containing individual elements, each of which is a 32-bit index value. The vector index register can be an XMM register (vm32x), a YMM register (vm32y), or a ZMM register (vm32z). In another embodiment, a memory operand of a SIB type with the form vm64{x,y,z} can identify a vector array of memory operands specified by using a memory addressing of the SIB type. In this example, the array of memory addresses is specified by using the following: a common base register, a constant scaling factor, and a vector index register containing individual elements, each of which is a 64-bit index value. The vector index register can be an XMM register (vm64x), a YMM register (vm64y), or a ZMM register (vm64z).

[0063] Figure 4 An embodiment of a method performed by a processor to process a fused multiply-accumulate instruction is illustrated.

[0064] At 401, an instruction is fetched. For example, a fused multiply-accumulate instruction is fetched. The fused multiply-accumulate instruction includes an opcode, and fields for a packed data source operand and a packed data destination operand as described in detail above. In some embodiments, the fused multiply-accumulate instruction includes a write mask operand. In some embodiments, the instruction is fetched from an instruction cache.

[0065] The fetched instruction is decoded at 403. For example, the fetched fused multiply-accumulate instruction is decoded by a decode circuit such as that described in detail herein.

[0066] At 405 , data values ​​associated with source operands of the decoded instruction are retrieved.

[0067] At 407, the decoded instruction is executed by execution circuitry (hardware) such as that described in detail herein. For a fused multiply-accumulate instruction, the execution will cause, for each packed data element location of the destination: a multiplication of M N-sized packed data elements from the first and second sources corresponding to the packed data element location of the third source, adding the results from the multiplications to the full-sized (relative to the N-sized packed data elements) packed data elements of the packed data element location of the third source, and storing the results of the addition(s) in the packed data element location destination corresponding to the packed data element location of the third source, where M equals the full-sized packed data element divided by N.

[0068] In some embodiments, the instruction is committed or retired at 409 .

[0069] Figure 5 An embodiment of a subset of the execution of fused multiply-accumulate is illustrated. In particular, this illustrates the execution circuitry of an iteration of one packed data element positioning of the destination. In this embodiment, the fused multiply-accumulate works on a signed source, where the accumulator is 2x the input data size. Figure 6 An embodiment of a pseudo code for implementing the instruction in hardware is illustrated.

[0070] The first signed source (source 1 501) and the second signed source (source 2 503) each have four packed data elements. Each of these packed data elements stores signed data, such as floating point data. The third signed source 509 (source 3) has two packed data elements, each of which stores signed data. The size of the first and second signed sources 501 and 503 is half the size of the third signed source 509. For example, the first and second signed sources 501 and 503 can have 32-bit packed data elements (e.g., single-precision floating point), and the third signed source 509 can have 64-bit packed data elements (e.g., double-precision floating point).

[0071] In this illustration, only the two most significant packed data element positions of the first and second signed sources 501 and 503 and the most significant packed data element position of the third signed source 509 are shown. Of course, other packed data element positions will also be processed.

[0072] As illustrated, the packed data elements are processed in pairs. For example, the data of the most significant packed data element location of the first and second signed sources 501 and 503 are multiplied by using multiplier circuit 505, and the data of the second most significant packed data element location from the first and second signed sources 501 and 503 are multiplied by using multiplier circuit 507. In certain embodiments, these multiplier circuits 505 and 507 are reused for other packed data element locations. In other embodiments, additional multiplier circuits are used to process the packed data elements in parallel. In some contexts, parallel execution is performed by using a track of the size of the third source 509 that is signed. The result of each multiplication is added by using addition circuit 511.

[0073] The result of the addition of the multiplication results is added to the data from the most significant packed data element position from signed source 3 509 (using either a different adder 513 or the same adder 511).

[0074] Finally, the result of the second addition is stored into signed destination 515 in a packed data element location corresponding to the packed data element location used from signed third source 509. In some embodiments, a write mask is applied to the store so that if the corresponding write mask (bit) is set, the store occurs, and if it is not set, the store does not occur.

[0075] Figure 7 An embodiment of a subset of the execution of fused multiply-accumulate is illustrated. In particular, this illustrates the execution circuitry of an iteration of one packed data element positioning of the destination. In this embodiment, the fused multiply-accumulate works on a signed source, where the accumulator is 2x the input data size. Figure 8 An embodiment of a pseudo code for implementing the instruction in hardware is illustrated.

[0076] The first signed source (source 1 701) and the second signed source (source 2 703) each have four packed data elements. Each of these packed data elements stores signed data, such as integer data. The third signed source 709 (source 3) has two packed data elements, each of which stores signed data. The size of the first and second signed sources 701 and 703 is half the size of the third signed source 709. For example, the first and second signed sources 701 and 703 can have 32-bit packed data elements (e.g., single-precision floating point), and the third signed source 709 can have 64-bit packed data elements (e.g., double-precision floating point).

[0077] In this illustration, only the two most significant packed data element positions of the first and second signed sources 701 and 703 and the most significant packed data element position of the third signed source 709 are shown. Of course, other packed data element positions will also be processed.

[0078] As illustrated, the packed data elements are processed in pairs. For example, the data of the most significant packed data element location of the first and second signed sources 701 and 703 are multiplied by using multiplier circuit 705, and the data of the second most significant packed data element location from the first and second signed sources 701 and 703 are multiplied by using multiplier circuit 707. In certain embodiments, these multiplier circuits 705 and 707 are reused for other packed data element locations. In other embodiments, additional multiplier circuits are used to process the packed data elements in parallel. In some contexts, parallel execution is performed by using a track of the size of the signed third source 709. The result of each multiplication is added to the signed third source 709 by using addition / saturation circuit 711.

[0079] Addition / saturation (accumulator) circuit 711 maintains the sign of the operands when addition results in a value that is too large. In particular, saturation evaluation occurs on infinite precision results between multi-way addition and writing to signed destination 715. When accumulator 711 is floating point and the input terms are integers, the sum of products and floating point accumulator input values ​​are converted to infinite precision values ​​(hundreds of digits of fixed point), addition of the multiplication result and the third input is performed, and single rounding to the actual accumulator type is performed.

[0080] The result of the addition and saturation check is stored into signed destination 715 in a packed data element location corresponding to the packed data element location used from signed third source 709. In some embodiments, a write mask is applied to the store so that if the corresponding write mask (bit) is set, the store occurs, and if it is not set, the store does not occur.

[0081] Fig. 9 An embodiment of a subset of the execution of a fused multiply-accumulate is illustrated. In particular, this illustrates the execution circuitry of an iteration of one packed data element positioning of the destination. In this embodiment, the fused multiply-accumulate works on signed sources as well as unsigned sources, where the accumulator is 4x the input data size. Fig.10 An embodiment of a pseudo code for implementing the instruction in hardware is illustrated.

[0082] The first signed source (source 1 901) and the second unsigned source (source 2 903) each have four packed data elements. Each of these packed data elements, such as floating point or integer data. The third signed source (source 3 915) has packed data elements in which signed data is stored. The size of the first and second sources 901 and 903 is one-fourth of the size of the third signed source 915. For example, the first and second sources 901 and 903 can have 16-bit packed data elements (e.g., words), and the third signed source 915 can have 64-bit packed data elements (e.g., double precision floating point or 64-bit integer).

[0083] In this illustration, the four most significant packed data element positions of the first and second sources 901 and 903 are shown, as well as the most significant packed data element position of the third signed source 915. Of course, other packed data element positions will also be processed, if any exist.

[0084] As illustrated, the packed data elements are processed in quadruplets. For example, data at the most significant packed data element positions of the first and second sources 901 and 903 are multiplied using multiplier circuit 907, data at the second most significant packed data element positions from the first and second sources 901 and 903 are multiplied using multiplier circuit 907, data at the third most significant packed data element positions from the first and second sources 901 and 903 are multiplied using multiplier circuit 909, and data at the least significant packed data element positions from the first and second sources 901 and 903 are multiplied using multiplier circuit 911. In some embodiments, the signed packed data elements of the first source 901 are sign extended and the unsigned packed data elements of the second source 903 are zero extended prior to multiplication.

[0085] In some embodiments, these multiplier circuits 905-911 are reused for other packed data element locations. In other embodiments, additional multiplier circuits are used so that the packed data elements are processed in parallel. In some contexts, parallel execution is performed by using a channel that is the size of the signed third source 915. The results of each multiplication are added by using an addition circuit 911.

[0086] The result of the addition of the multiplication results is added to the data at the most significant packed data element position from signed source 3 915 (by using either a different adder 913 or the same adder 911).

[0087] Finally, the result of the second addition is stored into signed destination 919 in a packed data element location corresponding to the packed data element location used from signed third source 909. In some embodiments, a write mask is applied to the store so that if the corresponding write mask (bit) is set, the store occurs, and if it is not set, the store does not occur.

[0088] Fig.11 An embodiment of a subset of the execution of a fused multiply-accumulate is illustrated. In particular, this illustrates the execution circuitry of an iteration of one packed data element positioning of the destination. In this embodiment, the fused multiply-accumulate works on signed sources as well as unsigned sources, where the accumulator is 4x the input data size. Fig.12 An embodiment of a pseudo code for implementing the instruction in hardware is illustrated.

[0089] The first signed source (source 1 1101) and the second unsigned source (source 2 1103) each have four packed data elements. Each of these packed data elements, such as floating point or integer data. The third signed source (source 3 1115) has packed data elements in which signed data is stored. The size of the first and second sources 1101 and 1103 is one-fourth of the size of the third signed source 1115. For example, the first and second sources 1101 and 1103 can have 16-bit packed data elements (e.g., words), and the third signed source 1115 can have 64-bit packed data elements (e.g., double precision floating point or 64-bit integer).

[0090] In this illustration, the four most significant packed data element positions of the first and second sources 1101 and 1103 are shown, as well as the most significant packed data element position of the third signed source 1115. Of course, other packed data element positions will also be processed, if any exist.

[0091] As illustrated, the packed data elements are processed in quadruplets. For example, data at the most significant packed data element positions of first and second sources 1101 and 1103 are multiplied using multiplier circuit 1107, data at the second most significant packed data element positions from first and second sources 1101 and 1103 are multiplied using multiplier circuit 1107, data at the third most significant packed data element positions from first and second sources 1101 and 1103 are multiplied using multiplier circuit 1109, and data at the least significant packed data element positions from first and second sources 1101 and 1103 are multiplied using multiplier circuit 1111. In some embodiments, the signed packed data elements of first source 1101 are sign extended and the unsigned packed data elements of second source 1103 are zero extended prior to multiplication.

[0092] In some embodiments, these multiplier circuits 1105-1111 are reused for other packed data element locations. In other embodiments, additional multiplier circuits are used so that the packed data elements are processed in parallel. In some contexts, parallel execution is performed by using a channel that is the size of the signed third source 1115. The result of the addition of the multiplication results is added to the data from the most significant packed data element location of the signed source 3 1115, which is added to the signed third source 1115 using the addition / saturation circuit 1113.

[0093] Addition / saturation (accumulator) circuit 1113 maintains the sign of the operand when an addition results in a value that is too large. In particular, saturation evaluation occurs on the infinite precision result between the multi-way addition and the write to destination 1115. When accumulator 1113 is floating point and the input terms are integers, the sum of products and the floating point accumulator input value are converted to infinite precision values ​​(hundreds of digits of fixed point), addition of the multiplication result and the third input is performed, and single rounding to the actual accumulator type is performed.

[0094] The result of the addition and saturation check is stored to signed destination 1119 in a packed data element location corresponding to the packed data element location used from signed third source 715. In some embodiments, a write mask is applied to the store so that if the corresponding write mask (bit) is set, the store occurs, and if it is not set, the store does not occur.

[0095] The following figures detail exemplary architectures and systems for implementing the above embodiments. In some embodiments, one or more of the hardware components and / or instructions described above are simulated as described below, or implemented as software modules.

[0096] An exemplary embodiment includes a processor including a decoder for decoding a single instruction having an opcode, a destination field representing a destination operand, and fields for first, second, and third packed data source operands, wherein packed data elements of the first and second packed data source operands have a first size different from a second size of packed data elements of a third packed data operand; a register file having a plurality of packed data registers including registers for source and destination operands; and execution circuitry for executing the decoded single instruction to perform, for each packed data element location of the destination operand: a multiplication of M number of N-sized packed data elements from the first and second packed data sources corresponding to a packed data element location of a third packed data source, adding results from the multiplications to full-sized packed data elements at packed data element locations of the third packed data source, and storing the addition results in a packed data element location destination corresponding to the packed data element location of the third packed data source, wherein M equals the full-sized packed data element divided by N.

[0097] In some embodiments, one or more of the following applies: the instruction defines a size of the packed data elements; the execution circuitry zero-extends the packed data elements of the second source and sign-extends the packed data elements of the first source before the multiplication; when the first size is half the second size, a first addition is performed on each of the multiplications and a second addition is performed on a result of the first addition and a result from a previous iteration; when the first size is half the second size, a single addition and a saturation check is performed on each of the multiplications and a result from a previous iteration; when the first size is one-quarter the second size, a first addition is performed on each of the multiplications and a second addition is performed on a result of the first addition and a result from a previous iteration; and / or when the first size is one-quarter the second size, a single addition and a saturation check is performed on each of the multiplications and a result from a previous iteration.

[0098] An exemplary embodiment includes a method for decoding a single instruction having an opcode, a destination field representing a destination operand, and fields for first, second, and third packed data source operands, wherein packed data elements of the first and second packed data source operands have a first size different from a second size of packed data elements of a third packed data operand; a register file having a plurality of packed data registers including registers for source and destination operands; and executing the decoded single instruction to perform, for each packed data element location of the destination operand: a multiplication of M number of N-sized packed data elements from the first and second packed data sources corresponding to a packed data element location of a third packed data source, adding results from the multiplications to full-sized packed data elements at packed data element locations of the third packed data source, and storing the addition results in a packed data element location destination corresponding to the packed data element location of the third packed data source, wherein M equals the full-sized packed data element divided by N.

[0099] In some embodiments, one or more of the following applies: the instruction defines a size of the packed data elements; the execution circuitry zero-extends the packed data elements of the second source and sign-extends the packed data elements of the first source before the multiplication; when the first size is half the second size, a first addition is performed on each of the multiplications and a second addition is performed on a result of the first addition and a result from a previous iteration; when the first size is half the second size, a single addition and a saturation check is performed on each of the multiplications and a result from a previous iteration; when the first size is one-quarter the second size, a first addition is performed on each of the multiplications and a second addition is performed on a result of the first addition and a result from a previous iteration; and / or when the first size is one-quarter the second size, a single addition and a saturation check is performed on each of the multiplications and a result from a previous iteration.

[0100] An exemplary embodiment includes a non-transitory machine-readable medium storing instructions that, when executed, cause a method for decoding a single instruction having an opcode, a destination field representing a destination operand, and fields for first, second, and third packed data source operands, wherein packed data elements of the first and second packed data source operands have a first size different from a second size of packed data elements of a third packed data operand; a register file having a plurality of packed data registers including registers for source and destination operands; and executing the decoded single instruction to perform, for each packed data element location of the destination operand: multiplications of M number of N-sized packed data elements from the first and second packed data sources corresponding to a packed data element location of a third packed data source, adding results from the multiplications to full-sized packed data elements at packed data element locations of the third packed data source, and storing the addition results in a packed data element location destination corresponding to a packed data element location of the third packed data source, wherein M equals the full-sized packed data element divided by N.

[0101] In some embodiments, one or more of the following applies: the instruction defines a size of the packed data elements; the execution circuitry zero-extends the packed data elements of the second source and sign-extends the packed data elements of the first source before the multiplication; when the first size is half the second size, a first addition is performed on each of the multiplications and a second addition is performed on a result of the first addition and a result from a previous iteration; when the first size is half the second size, a single addition and a saturation check is performed on each of the multiplications and a result from a previous iteration; when the first size is one-quarter the second size, a first addition is performed on each of the multiplications and a second addition is performed on a result of the first addition and a result from a previous iteration; and / or when the first size is one-quarter the second size, a single addition and a saturation check is performed on each of the multiplications and a result from a previous iteration.

[0102] An exemplary embodiment includes a system including a memory and a processor including a decoder for decoding a single instruction having an opcode, a destination field representing a destination operand, and fields for first, second, and third packed data source operands, wherein packed data elements of the first and second packed data source operands have a first size different from a second size of packed data elements of a third packed data operand; a register file having a plurality of packed data registers including registers for source and destination operands; and execution circuitry for executing the decoded single instruction to perform, for each packed data element location of the destination operand: multiplications of M number of N-sized packed data elements from the first and second packed data sources corresponding to a packed data element location of a third packed data source, adding results from the multiplications to full-sized packed data elements at packed data element locations of the third packed data source, and storing the addition results in a packed data element location destination corresponding to the packed data element location of the third packed data source, wherein M is equal to the full-sized packed data element divided by N.

[0103] In some embodiments, one or more of the following applies: the instruction defines a size of the packed data elements; the execution circuitry zero-extends the packed data elements of the second source and sign-extends the packed data elements of the first source before the multiplication; when the first size is half the second size, a first addition is performed on each of the multiplications and a second addition is performed on a result of the first addition and a result from a previous iteration; when the first size is half the second size, a single addition and a saturation check is performed on each of the multiplications and a result from a previous iteration; when the first size is one-quarter the second size, a first addition is performed on each of the multiplications and a second addition is performed on a result of the first addition and a result from a previous iteration; and / or when the first size is one-quarter the second size, a single addition and a saturation check is performed on each of the multiplications and a result from a previous iteration.

[0104] The embodiments of the (multiple) instructions described in detail above are embodied in the "universal vector friendly instruction format" described in detail below, and can be so embodied. In other embodiments, such a format is not utilized, and another instruction format is used, however, the following description of write mask registers, various data transformations (swizzle, broadcast, etc.), addressing, etc. is generally applicable to the description of the embodiments of the (multiple) instructions above. In addition, exemplary systems, architectures, and pipelines are described in detail below. The embodiments of the (multiple) instructions above can be executed on such systems, architectures, and pipelines, but are not limited to those described in detail.

[0105] Instruction Set

[0106] An instruction set may include one or more instruction formats. A given instruction format may define various fields (e.g., number of bits, position of bits) to specify, among other things, the operation to be performed (e.g., opcode) and the operands and / or other data fields (e.g., masks) on which the operation will be performed. Despite the definition of instruction templates (or subformats), some instruction formats are further decomposed. For example, an instruction template of a given instruction format may be defined to have different subsets of instruction format fields (the included fields are typically in the same order, but at least some have different bit positions because fewer fields are included) and / or defined to have given fields interpreted differently. Thus, each instruction of an ISA is expressed using a given instruction format (and if defined, with a given one of the instruction templates of the instruction format), and includes fields for specifying operations and operands. For example, an exemplary ADD (addition) instruction has a specific opcode and instruction format, the instruction format including an opcode field for specifying the opcode and an operand field for selecting operands (source 1 / destination and source 2); and an appearance of the ADD (addition) instruction in an instruction stream will have specific content in the operand field that selects a specific operand. A set of SIMD extensions known as Advanced Vector Extensions (AVX) (AVX1 and AVX2) and using a Vector Extensions (VEX) encoding scheme has been released and / or announced (e.g., see 64 and IA-32 Architectures Software Developer's Manual, September 2014; and see Advanced Vector Extensions Programming Reference, October 2014).

[0107] Example instruction format

[0108] Embodiments of the (multiple) instructions described herein may be embodied in different formats. In addition, exemplary systems, architectures, and pipelines are described in detail below. Embodiments of the (multiple) instructions may be executed on such systems, architectures, and pipelines, but are not limited to those described in detail.

[0109] Generic vector friendly instruction format

[0110] The vector friendly instruction format is an instruction format suitable for vector instructions (eg, there are certain fields specific to vector operations). Although embodiments are described in which both vector and scalar operations are supported by the vector friendly instruction format, alternative embodiments use only vector operations, the vector friendly instruction format.

[0111] Figures 13A-13B is a block diagram illustrating a general vector friendly instruction format and instruction templates thereof according to an embodiment of the present invention. Fig.13Ais a block diagram illustrating a general vector-friendly instruction format and a class A instruction template thereof according to an embodiment of the present invention; and Fig. 13B 1 is a block diagram illustrating a generic vector friendly instruction format and its class B instruction templates according to an embodiment of the present invention. In particular, a generic vector friendly instruction format 1300, for which class A and class B instruction templates are defined, both of which include non-memory access 1305 instruction templates and memory access 1320 instruction templates. The term "generic" in the context of the vector friendly instruction format means that the instruction format is not constrained to any specific instruction set.

[0112] Although embodiments of the present invention will be described, the vector friendly instruction format supports the following: 64-byte vector operand length (or size) with 32-bit (4-byte) or 64-bit (8-byte) data element width (or size) (and thus, a 64-byte vector includes 16 doubleword-sized elements or alternatively 8 quadword-sized elements); 64-byte vector operand length (or size) with 16-bit (2-byte) or 8-bit (1-byte) data element width (or size); 32-byte vector operand length (or size) with 32-bit (4-byte) or 64-bit (8-byte) data element width (or size); , 64-bit (8-byte), 16-bit (2-byte) or 8-bit (1-byte) data element width (or size); and 16-byte vector operand length (or size) with 32-bit (4-byte), 64-bit (8-byte), 16-bit (2-byte) or 8-bit (1-byte) data element width (or size); however, alternative embodiments may support more, fewer and / or different vector operand sizes (e.g., 256-byte vector operands) with more, fewer or different data element widths (e.g., 128-bit (16-byte) data element width).

[0113] Fig.13A The Class A instruction templates include: 1) within the non-memory access 1305 instruction template, there are shown a non-memory access, full round control type operation 1310 instruction template, and a non-memory access, data transformation type operation 1315 instruction template; and 2) within the memory access 1320 instruction template, there are shown a memory access, temporary 1325 instruction template, and a memory access, non-temporary 1330 instruction template. Fig. 13B The Class B instruction templates include: 1) within the non-memory access 1305 instruction template, there are shown a non-memory access, write mask control, partial round control type operation 1312 instruction template, and a non-memory access, write mask control, v size (vsize) type operation 1317 instruction template; and 2) within the memory access 1320 instruction template, there are shown a memory access, write mask control 1327 instruction template.

[0114] The general vector friendly instruction format 1300 includes the following: Figures 13A-13B The following fields are listed in the order shown in the figure.

[0115] Format field 1340 - The specific value in this field (the instruction format identifier value) uniquely identifies the vector friendly instruction format, and thus identifies an occurrence of an instruction in the vector friendly instruction format in an instruction stream. Thus, this field is optional in the sense that it is not required for instruction sets that only have the generic vector friendly instruction format.

[0116] Basic operation field 1342 - its content distinguishes different basic operations.

[0117] Register index field 1344 - its contents specify the location of the source and destination operands, whether they are in registers or in memory, either directly or through address generation. These include a sufficient number of bits to select N registers from a PxQ (e.g., 32x512, 16x128, 32x1024, 64x1024) register file. Although in one embodiment N may be up to three sources and one destination register, alternative embodiments may support more or fewer source and destination registers (e.g., may support up to two sources, one of which also serves as a destination; may support up to three sources, one of which also serves as a destination; may support up to two sources and one destination).

[0118] Modifier field 1346 - its contents distinguish between occurrences of instructions in the generic vector instruction format that specify memory access and those that do not; that is, between non-memory access 1305 instruction templates and memory access 1320 instruction templates. Memory access operations read and / or write to the memory hierarchy (in some cases by specifying source and / or destination addresses using values ​​in registers), while non-memory access operations do not (e.g., the source and destination are registers). While in one embodiment this field also selects between three different ways to perform memory address calculations, alternative embodiments may support more, fewer, or different ways to perform memory address calculations.

[0119] Augmented operation field 1350 - its content distinguishes which of the various different operations will be performed in addition to the basic operation. This field is context specific. In one embodiment of the invention, this field is divided into a class field 1368, an alpha field 1352, and a beta field 1354. Augmented operation field 1350 allows a common group of operations to be performed in a single instruction instead of 2, 3, or 4 instructions.

[0120] Scale field 1360 - its content allows scaling of the contents of the index field for memory address generation (e.g., for use with "2 缩放 *index+base(2 scale *index+base)”).

[0121] Displacement field 1362A - its contents are used as part of memory address generation (e.g., for use with "2 缩放 * index + base + displacement (2 scale *index+base+displacement)”).

[0122] Displacement Factor field 1362B (note that the concatenation of displacement field 1362A directly above displacement factor field 1362B indicates that one or the other is used) - its contents are used as part of the address generation; it specifies the displacement factor that will be scaled by the size (N) of the memory access - where N is the number of bytes in the memory access (e.g., for use with "2 缩放 * index + base + scaled displacement (2 scale The redundant low order bits are ignored and therefore, the contents of the displacement factor field are multiplied by the total size of the memory operand (N) to generate the final displacement to be used in calculating the effective address. The value of N is determined by the processor hardware at runtime based on the full opcode field 1374 (described later herein) and the data manipulation field 1354C. The displacement field 1362A and the displacement factor field 1362B are optional in the sense that they are not used for the non-memory access 1305 instruction templates and / or different embodiments may implement only one of the two or neither of the two.

[0123] Data element width field 1364 - its contents distinguish which of multiple data element widths will be used (in some embodiments for all instructions; in other embodiments only for some of the instructions). This field is optional in the sense that it is not needed if only one data element width is supported and / or if data element width is supported by using some aspect of the opcode.

[0124] Write mask field 1370 - its content controls whether the data element position in the destination vector operand reflects the result of the base and augmentation operations on a per-data element position basis. Class A instruction templates support merge-writemasking, while class B instruction templates support both merge- and zero-writemasking. When merging, the vector mask allows any set of elements in the destination to be protected from updates during the execution of any operation (specified by the base and augmentation operations); in another embodiment, the old value of each element of the destination where the corresponding mask bit has a 0 is retained. In contrast, when zeroing, the vector mask allows any set of elements in the destination to be zeroed during the execution of any operation (specified by the base and augmentation operations); in one embodiment, when the corresponding mask bit has a 0 value, the elements of the destination are set to 0. A subset of this functionality is the ability to control the vector length (i.e., the elements modified, the span from the first to the last) of the operation being performed; however, it is not necessary for the modified elements to be sequential. Thus, the write mask field 1370 allows partial vector operations, including loads, stores, arithmetic, logical, etc. Although embodiments of the present invention have been described in which the contents of write mask field 1370 select one of a plurality of write mask registers containing a write mask to be used (and thus the contents of write mask field 1370 indirectly identify the masking to be performed), alternative embodiments alternatively or additionally allow the contents of mask write field 1370 to directly specify the masking to be performed.

[0125] Immediate field 1372 - its content allows specification of an immediate. This field is optional in the sense that it is not present in implementations that do not support the generic vector friendly format and it is not present in instructions that do not use immediates.

[0126] Class field 1368 - its content distinguishes between instructions of different classes. Fig.13A -B, the contents of this field select between class A and class B instructions. Fig.13A -B, the rounded corner square is used to indicate the presence of a specific value in the field (for example, Fig.13A - Class A 1368A and Class B 1368B for class field 1368 in B).

[0127] Class A instruction template

[0128] In the case of the class A non-memory access 1305 instruction templates, the alpha field 1352 is interpreted as the RS field 1352A, whose contents distinguish which of the different augmentation operation types are to be performed (e.g., rounding 1352A.1 and data transformation 1352A.2 are specified for the non-memory access, rounding type operation 1310 and non-memory access, data transformation type operation 1315 instruction templates, respectively), while the beta field 1354 distinguishes which of the specified types of operations are to be performed. In the non-memory access 1305 instruction templates, the scale field 1360, displacement field 1362A, and displacement scale field 1362B are not present.

[0129] Non-memory access instruction templates - full rounding control type operations

[0130] In the non-memory access full rounding control type operation 1310 instruction template, the beta field 1354 is interpreted as a rounding control field 1354A, whose (multiple) contents provide static rounding. Although in the described embodiment of the present invention, the rounding control field 1354A includes a suppress all floating point exceptions (SAE) field 1356 and a rounding operation control field 1358, alternative embodiments may support, may encode these two concepts into the same field, or have only one or the other of these concepts / fields (e.g., may only have a rounding operation control field 1358).

[0131] SAE field 1356 - its content distinguishes whether to disable exception event reporting; when the content of SAE field 1356 indicates that suppression is enabled, the given instruction does not report any kind of floating point exception flags and does not trigger any floating point exception handler.

[0132] Rounding operation control field 1358 - its content distinguishes which of the group of rounding operations is to be performed (e.g., round up, round down, round to zero, and round to nearest). Thus, the rounding operation control field 1358 allows the rounding mode to be changed on a per-instruction basis. In one embodiment of the invention in which the processor includes a control register for specifying the rounding mode, the content of the rounding operation control field 1350 overwrites the register value.

[0133] Non-memory access instruction templates - data transformation type operations

[0134] In a non-memory access data transformation type operation 1315 instruction template, the beta field 1354 is interpreted as a data transformation field 1354B, whose content distinguishes which of multiple data transformations is to be performed (eg, no data transformation, swizzling, broadcasting).

[0135] In the case of the class A memory access 1320 instruction template, the alpha field 1352 is interpreted as an eviction hint field 1352B, the contents of which distinguish which of the eviction hints (in Fig.13A 2), and the beta field 1354 is interpreted as a data manipulation field 1354C, the contents of which distinguish which of a plurality of data manipulation operations (also known as primitives) is to be performed (e.g., no manipulation; broadcast; upcast of the source; and downcast of the destination). The memory access 1320 instruction template includes a scale field 1360, and optionally a displacement field 1362A or a displacement scale field 1362B.

[0136] Vector memory instructions perform vector loads from memory and vector stores to memory with conversion support. Like conventional vector instructions, vector memory instructions transfer data from / to memory in a per-data-element manner, where the actual elements transferred are indicated by the contents of the vector mask selected as the writemask.

[0137] Memory Access Instruction Templates - Temporary

[0138] Temporary data is data that is likely to be reused quickly enough to benefit from caching. However, this is a hint, and different processors can implement it in different ways, including ignoring the hint completely.

[0139] Memory Access Instruction Templates - Non-Temporal

[0140] Non-temporal data is data that is unlikely to be reused quickly enough to benefit from being cached in the L1 cache, and should be given priority for eviction. However, this is a hint, and different processors may implement it in different ways, including ignoring the hint entirely.

[0141] Class B instruction template

[0142] In the case of class B instruction templates, the alpha field 1352 is interpreted as a write mask control (Z) field 1352C, the content of which distinguishes whether the write masking controlled by the write mask field 1370 should be merging or zeroing.

[0143] In the case of the class B non-memory access 1305 instruction templates, a portion of the beta field 1354 is interpreted as an RL field 1357A, the contents of which distinguish which of the different augmentation operation types is to be performed (e.g., rounding 1357A.1 and vector length (VSIZE) 1357A.2 are specified for the non-memory access, write mask control, partial rounding control type operation 1312 instruction template and the non-memory access, write mask control, V-size (VSIZE) type operation 1317 instruction template, respectively), while the remainder of the beta field 1354 distinguishes which of the specified types of operations is to be performed. In the non-memory access 1305 instruction templates, the scale field 1360, displacement field 1362A, and displacement scale field 1362B are not present.

[0144] In a non-memory access, write mask control, partial rounding control type operation 1310 instruction template, the remainder of the beta field 1354 is interpreted as the rounding operation field 1359A and exception event reporting is disabled (the given instruction does not report any kind of floating point exception flags and does not raise any floating point exception handler).

[0145] Rounding operation control field 1359A - Just as with the rounding operation control field 1358, its content distinguishes which of the group of rounding operations is to be performed (e.g., round up, round down, round to zero, and round to nearest). Thus, the rounding operation control field 1359A allows the rounding mode to be changed on a per-instruction basis. In one embodiment of the invention in which the processor includes a control register for specifying the rounding mode, the content of the rounding operation control field 1350 overwrites the register value.

[0146] In a non-memory access, write mask control, V size (VSIZE) type operation 1317 instruction template, the remainder of the beta field 1354 is interpreted as a vector length field 1359B, the contents of which distinguish which of multiple data vector lengths will be performed on (e.g., 128, 256, or 512 bytes).

[0147] In the case of a class B memory access 1320 instruction template, a portion of the beta field 1354 is interpreted as a broadcast field 1357B, the contents of which distinguish whether a broadcast type of data manipulation operation is to be performed, and the remainder of the beta field 1354 is interpreted as a vector length field 1359B. The memory access 1320 instruction template includes a scale field 1360, and optionally a displacement field 1362A or a displacement scale field 1362B.

[0148] With respect to the generic vector friendly instruction format 1300, a full opcode field 1374 is shown that includes the format field 1340, the base operation field 1342, and the data element width field 1364. While one embodiment is shown where the full opcode field 1374 includes all of these fields, in embodiments that do not support all of them, the full opcode field 1374 includes less than all of these fields. The full opcode field 1374 provides an operation code (opcode).

[0149] The augmentation operation field 1350, the data element width field 1364, and the write mask field 1370 allow these features to be specified on a per-instruction basis in the generic vector friendly instruction format.

[0150] The combination of the write mask field and the data element width field creates a type instruction because they allow the mask to be applied based on different data element widths.

[0151] The various instruction templates present in class A and class B are beneficial in different situations. In some embodiments of the present invention, different processors or different cores within a processor may support only class A, only class B, or both classes. For example, a high-performance general-purpose out-of-order core intended for general-purpose computing may support only class B, a core intended primarily for graphics and / or scientific (throughput) computing may support only class A, and a core intended for both may support both (of course, a core with some mixture of templates and instructions from both classes, but not all templates and instructions from both classes, is within the scope of the present invention). Moreover, a single processor may include multiple cores, all of which support the same class or different cores of which support different classes. For example, in a processor with separate graphics and general-purpose cores, one of the graphics cores intended primarily for graphics and / or scientific computing may support only class A, while one or more of the general-purpose cores may be a high-performance general-purpose core with out-of-order execution and register renaming intended for general-purpose computing that supports only class B. Another processor without a separate graphics core may include one or more general-purpose in-order or out-of-order cores that support both class A and class B. Of course, in different embodiments of the present invention, features from one class may also be implemented in another class. A program written in a high-level language will be placed (e.g., compiled just in time or statically) in a variety of different executable forms, including: 1) a form that has only instructions of the class(es) supported by the target processor for execution; or 2) a form that has replaceable routines written using different combinations of instructions of all classes and with control flow code that selects the routine to execute based on the instructions supported by the processor currently executing the code.

[0152] Exemplary specific vector friendly instruction format

[0153] Fig.14A is a block diagram illustrating an exemplary specific vector friendly instruction format according to an embodiment of the present invention. Fig.14A A specific vector friendly instruction format 1400 is shown, which is specific in the sense that it specifies the location, size, interpretation, and order of fields, as well as values ​​for some of those fields. The specific vector friendly instruction format 1400 can be used to extend the x86 instruction set, and thus some of the fields are similar or identical to those used in the existing x86 instruction set and its extensions (e.g., AVX). The format remains consistent with the prefix encoding field, real opcode byte field, MOD R / M field, SIB field, displacement field, and immediate field of the existing x86 instruction set with extensions. Fig.14A The fields of are mapped to the fields from Figure 13 therein.

[0154] It should be understood that although embodiments of the present invention are described with reference to the specific vector friendly instruction format 1400 in the context of the generic vector friendly instruction format 1300 for illustrative purposes, the present invention is not limited to the specific vector friendly instruction format 1400, except where required. For example, the generic vector friendly instruction format 1300 contemplates various possible sizes for various fields, while the specific vector friendly instruction format 1400 is illustrated as having fields of specific sizes. As a specific example, although the data element width field 1364 is illustrated as a one-bit field in the specific vector friendly instruction format 1400, the present invention is not so limited (that is, the generic vector friendly instruction format 1300 contemplates other sizes of the data element width field 1364).

[0155] The general vector friendly instruction format 1300 includes the following: Fig.14A The following fields are listed in the order shown in the figure.

[0156] EVEX prefix (bytes 0-3) 1402 - encoded in four-byte form.

[0157] Format field 1340 (EVEX byte 0, bits [7:0]) - The first byte (EVEX byte 0) is the format field 1340 and it contains 0x62 (a unique value used to distinguish the vector friendly instruction format in one embodiment of the invention).

[0158] The second through fourth bytes (EVEX bytes 1-3) include a number of bit fields that provide specific capabilities.

[0159] REX field 1405 (EVEX byte 1, bits [7-5]) - includes the EVEX.R bit field (EVEX byte 1, bits [7]-R), the EVEX.X bit field (EVEX byte 1, bits [6]-X), and 1357BEX byte 1, bits [5]-B). The EVEX.R, EVEX.X, and EVEX.B bit fields provide the same functionality as the corresponding VEX bit fields and are encoded using multiple ones' complement form, i.e., ZMM0 is encoded as 1111B and ZMM15 is encoded as 0000B. The other fields of the instruction encode the lower three bits of the register index, as known in the art (rrr, xxx, and bbb), such that Rrrr, Xxxx, and Bbbb can be formed by adding EVEX.R, EVEX.X, and EVEX.B.

[0160] REX' Field 1310 - This is the first portion of the REX' field 1310 and is the EVEX.R' bit field (EVEX byte 1, bit [4] - R') which is used to encode the upper 16 or lower 16 of the extended 32 register set. In one embodiment of the invention, this bit, along with other bits as indicated below, are stored in a bit-reversed format to distinguish (in the well-known x86 32-bit mode) from the BOUND instruction, which actually has an opcode byte of 62, but does not accept a value of 11 in the MOD field in the MOD R / M field (described below); alternative embodiments of the invention do not store this and the other indicated bits below in an inverted format. The value 1 is used to encode the lower 16 registers. In other words, R'Rrrr is formed by combining EVEX.R', EVEX.R, and the other RRRs from the other fields.

[0161] Opcode Map field 1415 (EVEX byte 1, bits [3:0] - mmmm) - its contents encode the implied leading opcode byte (0F, 0F 38, or 0F 3).

[0162] Data element width field 1364 (EVEX byte 2, bit [7] - W) - This is represented by the notation EVEX.W. EVEX.W is used to define the granularity (size) of the data type (32-bit data element or 64-bit data element).

[0163] EVEX.vvvv 1420 (EVEX byte 2, bits [6:3] - vvvv) - The role of EVEX.vvvv may include the following: 1) EVEX.vvvv encodes the first source register operand specified in inverted (multiple ones complement) form and is valid for instructions with 2 or more source operands; 2) EVEX.vvvv encodes the destination register operand specified in multiple ones complement form for certain vector shifts; or 3) EVEX.vvvv does not encode any operand, the field is reserved and should contain 1111b. Thus, EVEX.vvvv field 1420 encodes the 4 low order bits of the first source register specifier stored in inverted (multiple ones complement) form. Depending on the instruction, additional different EVEX bit fields are used to extend the specifier size to 32 registers.

[0164] EVEX.U 1368 Class field (EVEX byte 2, bit [2] - U) - if EVEX.U = 0, it indicates Class A or EVEX.U0; if EVEX.U = 1, it indicates Class B or EVEX.U1.

[0165] Prefix encoding field 1425 (EVEX byte 2, bits [1:0]-pp) - provides additional bits for the basic operation field. In addition to providing support for legacy SSE instructions in EVEX prefix format, this has the benefit of making the SIMD prefix compact (the EVEX prefix only requires 2 bits, rather than requiring a byte to express the SIMD prefix). In one embodiment, to support legacy SSE instructions that use SIMD prefixes (66H, F2H, F3H) in both legacy and EVEX prefix formats, these legacy SIMD prefixes are encoded into the SIMD prefix encoding field; and expanded into the legacy SIMD prefix at runtime and then provided to the decoder's PLA (so the PLA can execute both legacy and EVEX formats of these legacy instructions without modification). Although newer instructions can use the contents of the EVEX prefix encoding field directly as an opcode extension, some embodiments expand in a similar manner for consistency, but allow different meanings to be specified by these legacy SIMD prefixes. An alternative embodiment may redesign the PLA to support 2-bit SIMD prefix encoding, and thus not require expansion.

[0166] Alpha field 1352 (EVEX byte 3, bit [7] - EH; also known as EVEX.EH, EVEX.rs, EVEX.RL, EVEX.WriteMaskControl, and EVEX.N; also illustrated using α) - As previously described, this field is context specific.

[0167] Beta field 1354 (EVEX byte 3, bits [6:4] - SSS, also known as EVEX.s 2-0 ,EVEX.r 2-0 , EVEX.rr1, EVEX.LL0, EVEX.LLB; also illustrated using βββ) - As previously described, this field is context-specific.

[0168] REX' field 1310 - This is the remainder of the REX' field and is the EVEX.V' bit field (EVEX byte 3, bit [3] - V') which can be used to encode the upper 16 or lower 16 of the extended 32 register set. This bit is stored in a bit-reversed format. A value of 1 is used to encode the lower 16 registers. In other words, V'VVVV is formed by combining EVEX.V', EVEX.vvvv.

[0169] Write mask field 1370 (EVEX byte 3, bits [2:0] - kkk) - its contents specify the index of a register in the write mask register, as previously described. In one embodiment of the invention, the particular value EVEX.kkk = 000 has a special appearance, meaning that no write mask is used for the particular instruction (this can be implemented in a variety of ways, including using a write mask that is hardwired to all ones or hardware that bypasses the masking hardware).

[0170] The real opcode field 1430 (byte 4) is also known as the opcode byte. The portion of the opcode is specified in this field.

[0171] The MOD R / M field 1440 (byte 5) includes a MOD field 1442, a Reg field 1444, and an R / M field 1446. As previously described, the content of the MOD field 1442 distinguishes between memory access and non-memory access operations. The role of the Reg field 1444 can be summarized into two cases: encoding a destination register operand or a source register operand, or being treated as an opcode extension and not used to encode any instruction operand. The role of the R / M field 1446 may include the following: encoding an instruction operand that references a memory address, or encoding a destination register operand or a source register operand.

[0172] Scale, Index, Base (SIB) Byte (Byte 6) - As previously described, the contents of the scale field 1350 are used for memory address generation. SIB.xxx 1454 and SIB.bbb 1456 - The contents of these fields have been mentioned previously with respect to register indexes Xxxx and Bbbb.

[0173] Displacement field 1362A (bytes 7-10) - When the MOD field 1442 contains 10, bytes 7-10 are the displacement field 1362A and it functions the same as the legacy 32-bit displacement (disp32) and functions at byte granularity.

[0174] Displacement Factor Field 1362B (Byte 7) - When the MOD field 1442 contains 01, byte 7 is the displacement factor field 1362B. The location of this field is the same as the location of the legacy x86 instruction set 8-bit displacement (disp8) that operates at byte granularity. Because disp8 is sign-extended, it can only address between -128 and 127 byte offsets; in terms of a 64-byte cache line, disp8 uses 8 bits that can be set to only four real useful values ​​-128, -64, 0, and 64; because a larger range is usually required, disp32 is used; however, disp32 requires 4 bytes. In contrast to disp8 and disp32, the displacement factor field 1362B is a reinterpretation of disp8; when the displacement factor field 1362B is used, the actual displacement is determined by multiplying the contents of the displacement factor field by the size (N) of the memory operand access. This type of displacement is referred to as disp8*N. This reduces the average instruction length (a single byte for displacement but with a much larger range). Such compressed displacement is based on the following assumption: namely, the effective displacement is a multiple of the granularity of the memory access, and therefore, the redundant low-order bits of the address offset do not need to be encoded. In other words, the displacement factor field 1362B replaces the displacement of the legacy x86 instruction set 8 bits. Thus, the displacement factor field 1362B is encoded in the same manner as the displacement of the x86 instruction set 8 bits (therefore, there is no change in the ModRM / SIB encoding rules), with the only exception that disp8 is overloaded to disp8*N. In other words, there is no change in the encoding rules or encoding length, but only in the interpretation of the displacement value by the hardware (which needs to scale the displacement by the size of the memory operand to obtain the address offset in bytes). The immediate field 1372 operates as previously described.

[0175] Full opcode field

[0176] Fig. 14B is a block diagram illustrating the fields of a particular vector friendly instruction format 1400 that make up a full opcode field 1374 according to one embodiment of the present invention. In particular, the full opcode field 1374 includes a format field 1340, a basic operation field 1342, and a data element width (W) field 1364. The basic operation field 1342 includes a prefix encoding field 1425, an opcode map field 1415, and a real opcode field 1430.

[0177] Register Index Field

[0178] Fig. 14C is a block diagram illustrating the fields of a particular vector friendly instruction format 1400 that make up a register index field 1344 according to one embodiment of the present invention. In particular, the register index field 1344 includes a REX field 1405, a REX' field 1410, a MODR / M.reg field 1444, a MODR / Mr / m field 1446, a VVVV field 1420, a xxx field 1454, and a bbb field 1456.

[0179] Augmentation Operation Field

[0180] Fig.14D is a block diagram illustrating the fields of a particular vector friendly instruction format 1400 that make up the augment operation field 1350 according to one embodiment of the present invention. When the class (U) field 1368 contains 0, it indicates EVEX.U0 (class A 1368A); when it contains 1, it indicates EVEX.U1 (class B 1368B). When U=0 and the MOD field 1442 contains 11 (indicating a non-memory access operation), the alpha field 1352 (EVEX byte 3, bit [7] - EH) is interpreted as the rs field 1352A. When the rs field 1352A contains a 1 (round 1352A.1), the beta field 1354 (EVEX byte 3, bits [6:4] - SSS) is interpreted as the round control field 1354A. The round control field 1354A includes a one-bit SAE field 1356 and a two-bit round operation field 1358. When the rs field 1352A contains a 0 (data transformation 1352A.2), the beta field 1354 (EVEX byte 3, bits [6:4]-SSS) is interpreted as a three-bit data transformation field 1354B. When U=0 and the MOD field 1442 contains 00, 01, or 10 (indicating a memory access operation), the alpha field 1352 (EVEX byte 3, bits [7]-EH) is interpreted as an eviction hint (EH) field 1352B, and the beta field 1354 (EVEX byte 3, bits [6:4]-SSS) is interpreted as a three-bit data manipulation field 1354C.

[0181] When U=1, the alpha field 1352 (EVEX byte 3, bit [7]-EH) is interpreted as the write mask control (Z) field 1352C. When U=1 and the MOD field 1442 contains 11 (indicating a non-memory access operation), part of the beta field 1354 (EVEX byte 3, bit [4]-S0) is interpreted as the RL field 1357A; when it contains a 1 (round 1357A.1), the remainder of the beta field 1354 (EVEX byte 3, bits [6-5]-S0) is interpreted as the RL field 1357A.2-1 ) is interpreted as the rounding operation field 1359A, and when the RL field 1357A contains a zero (VSIZE 1357.A2), the remainder of the beta field 1354 (EVEX byte 3, bits [6-5] - S 2-1 ) is interpreted as the vector length field 1359B (EVEX byte 3, bits [6-5] - L 1-0 When U=1 and the MOD field 1442 contains 00, 01, or 10 (indicating a memory access operation), the beta field 1354 (EVEX byte 3, bits [6:4] - SSS) is interpreted as the vector length field 1359B (EVEX byte 3, bits [6-5] - L 1-0 ) and broadcast field 1357B (EVEX byte 3, bit [4]-B).

[0182] Exemplary Register Architecture

[0183] Fig.15 1 is a block diagram of a register architecture 1500 according to one embodiment of the present invention. In the illustrated embodiment, there are 32 vector registers 1510 that are 512 bits wide; these registers are referenced as zmm0 through zmm31. The lower order 256 bits of the lower 16 zmm registers are overlaid on registers ymm0-16. The lower order 128 bits of the lower 16 zmm registers (the lower order 128 bits of the ymm registers) are overlaid on registers xmm0-15. Specific vector friendly instruction format

[0184] Equation 1400 operates on these overlaid register files as illustrated in the following table.

[0185]

[0186] In other words, the vector length field 1359B selects between a maximum length and one or more other shorter lengths, where each such shorter length is half as long as the previous length; and instruction templates without the vector length field 1359B operate on the maximum vector length. In addition, in one embodiment, the class B instruction templates of the specific vector friendly instruction format 1400 operate on packed or scalar single / double precision floating point data and packed or scalar integer data. Scalar operations are operations performed on the lowest order data element locations in the zmm / ymm / xmm registers; either the higher order data element locations are made the same as they were before the instruction, or are zeroed, depending on the embodiment.

[0187] Write mask registers 1515 - In the illustrated embodiment, there are 8 write mask registers (k0 through k7), each 64 bits in size. In an alternative embodiment, write mask registers 1515 are 16 bits in size. As previously described, in one embodiment of the invention, vector mask register k0 cannot be used as a write mask; while the encoding that would normally indicate k0 is used for write masking, it selects a hardwired write mask of 0xFFFF, which effectively disables write masking for that instruction.

[0188] General purpose registers 1525 - In the illustrated embodiment, there are sixteen 64-bit general purpose registers that are used along with existing x86 addressing modes to address memory operands. These registers are referenced by the names RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 through R15.

[0189] Scalar floating point stack register file (x87 stack) 1545, on which the MMX packed integer flat register file 1550 is aliased - in the illustrated embodiment, the x87 stack is an eight-element stack that is used to perform scalar floating point operations on 32 / 64 / 80-bit floating point data using the x87 instruction set extension; while the MMX registers are used to perform operations on 64-bit packed integer data, as well as to hold operands for some operations performed between MMX and XMM registers.

[0190] Alternative embodiments of the present invention may use wider or narrower registers. In addition, alternative embodiments of the present invention may use more, fewer or different register files and registers.

[0191] Exemplary Core Architectures, Processors, and Computer Architectures

[0192] Processor cores may be implemented in different ways, for different purposes, and in different processors. For example, implementations of such cores may include: 1) a general-purpose in-order core intended for general-purpose computing; 2) a high-performance general-purpose out-of-order core intended for general-purpose computing; 3) a specialized core intended primarily for graphics and / or scientific (throughput) computing. Different processor implementations may include: 1) a CPU including one or more general-purpose in-order cores intended for general-purpose computing and / or one or more general-purpose out-of-order cores intended for general-purpose computing; and 2) a coprocessor including one or more specialized cores intended primarily for graphics and / or scientific (throughput). Such different processors lead to different computer system architectures, which may include: 1) a coprocessor on a separate chip from the CPU; 2) a coprocessor on a separate die in the same package as the CPU; 3) a coprocessor on the same die as the CPU (in which case such coprocessors are sometimes referred to as dedicated logic, such as integrated graphics and / or scientific (throughput) logic, or as dedicated cores); and 4) a system on a chip that may include the CPU (sometimes referred to as application core(s) or application processor(s), the above coprocessor, and additional functionality on the same die. An exemplary core architecture is described next, followed by a description of an exemplary processor and computer architecture.

[0193] Exemplary Core Architecture

[0194] In-order and out-of-order core block diagram

[0195] Fig.16A is a block diagram illustrating both an exemplary in-order pipeline and an exemplary register renaming, out-of-order issue / execution pipeline according to embodiments of the present invention. Fig. 16B is a block diagram illustrating exemplary embodiments of both an in-order architecture core and an exemplary register renaming, out-of-order issue / execution architecture core to be included in a processor in accordance with embodiments of the present invention. Fig.16A The solid line boxes in -B illustrate the in-order pipeline and in-order core, while the optional addition of dashed line boxes illustrates the register renaming, out-of-order issue / execution pipeline and core. Considering that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.

[0196] exist Fig.16A , the processor pipeline 1600 includes a fetch stage 1602, a length decode stage 1604, a decode stage 1606, an allocation stage 1608, a rename stage 1610, a schedule (also known as dispatch or issue) stage 1612, a register read / memory read stage 1614, an execute stage 1616, a write back / memory write stage 1618, an exception handling stage 1622, and a commit stage 1624.

[0197] Fig. 16B A processor core 1690 is shown, the processor core 1690 includes a front end unit 1630, the front end unit 1630 is coupled to an execution engine unit 1650, and the front end unit 1630 and the execution engine unit 1650 are both coupled to a memory unit 1670. The core 1690 can be a reduced instruction set computing (RISC) core, a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or replaceable core type. As another option, the core 1690 can be a special-purpose core, such as, for example, a network or communication core, a compression engine, a coprocessor core, a general-purpose computing graphics processing unit (GPGPU) core, a graphics core, etc.

[0198] The front end unit 1630 includes a branch prediction unit 1632, which is coupled to an instruction cache unit 1634, which is coupled to an instruction translation lookaside buffer (TLB) 1636, which is coupled to an instruction fetch unit 1638, which is coupled to a decode unit 1640. The decode unit 1640 (or decoder) can decode instructions and generate one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals as output, which are decoded from, or otherwise reflect, or are derived from the original instructions. The decode unit 1640 can be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLA), microcode read-only memories (ROMs), and the like. In one embodiment, the core 1690 includes a microcode ROM or other medium that stores microcode for certain macroinstructions (e.g., in the decode unit 1640 or otherwise within the front end unit 1630). Decode unit 1640 is coupled to rename / allocator unit 1652 in execution engine unit 1650 .

[0199] The execution engine unit 1650 includes a rename / allocator unit 1652 coupled to a retirement unit 1654 and a set of one or more scheduler units 1656. The scheduler unit(s) 1656 represent any number of different schedulers, including reservation stations, central instruction windows, and the like. The scheduler unit(s) 1656 are coupled to physical register file(s) units 1658. Each of the physical register file(s) units 1658 represents one or more physical register files, different ones of which store one or more different data types, such as scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points, status (e.g., an instruction pointer that is the address of the next instruction to be executed), and the like. In one embodiment, the physical register file(s) units 1658 include a vector register unit, a write mask register unit, and a scalar register unit. These register units may provide architectural vector registers, vector mask registers, and general purpose registers. Physical register(s) file(s) unit(s) 1658 are overlapped by retirement unit(s) 1654 to illustrate various ways in which register renaming and out-of-order execution may be implemented (e.g., by using reorder buffer(s) and retirement register file(s); by using future file(s), history buffer(s), and retirement register file(s); by using register maps and register pools; etc.). Retirement unit(s) 1654 and physical register(s) file(s) unit(s) 1658 are coupled to execution cluster(s) 1660. Execution cluster(s) 1660 include a set of one or more execution units 1662 and a set of one or more memory access units 1664. Execution units 1662 may perform various operations (e.g., shifts, additions, subtractions, multiplications) and may perform on various types of data (e.g., scalar floating point, packed integer, packed floating point, vector integer, vector floating point). While some embodiments may include multiple execution units dedicated to specific functions or sets of functions, other embodiments may include only one execution unit or multiple execution units that all perform all functions. Scheduler unit(s) 1656, physical register file(s) units 1658, and execution cluster(s) 1660 are shown as possibly plural because certain embodiments create separate pipelines for certain types of data / operations (e.g., scalar integer pipelines, scalar floating point / packed integer / packed floating point / vector integer / vector floating point pipelines, and / or memory access pipelines, each of which has its own scheduler unit, physical register file(s) units, and / or execution clusters - and in the case of separate memory access pipelines, certain embodiments are implemented where only the execution cluster for that pipeline has memory access unit(s) 1664).It should also be understood that in cases where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution and the rest in-order.

[0200] The set of memory access units 1664 is coupled to a memory unit 1670, which includes a data TLB unit 1672, which is coupled to a data cache unit 1674, which is coupled to a level 2 (L2) cache unit 1676. In an exemplary embodiment, the memory access unit 1664 may include a load unit, a store address unit, and a store data unit, each of which is coupled to the data TLB unit 1672 in the memory unit 1670. The instruction cache unit 1634 is further coupled to a level 2 (L2) cache unit 1676 in the memory unit 1670. The L2 cache unit 1676 is coupled to one or more other levels of cache and ultimately to main memory.

[0201] As an example, the exemplary register renaming, out-of-order issue / execution core architecture may implement pipeline 1600 as follows: 1) instruction fetch 1638 performs fetch and length decode stages 1602 and 1604; 2) decode unit 1640 performs decode stage 1606; 3) rename / allocator unit 1652 performs allocate stage 1608 and rename stage 1610; 4) (multiple) scheduler unit 1656 performs schedule stage 1612; 5) physical register(s) file(s) unit 1658 and memory unit 1670 perform register read / memory read stage 1614; execution cluster 1660 performs execute stage 1616; 6) memory unit 1670 and physical register(s) file(s) unit 1658 perform write back / memory write stage 1618; 7) various units may be involved in exception handling stage 1622; and 8) retirement unit 1654 and physical register(s) file(s) unit 1658 perform commit stage 1624.

[0202] Core 1690 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions that have been added in the case of newer versions); the MIPS instruction set from MIPS Technologies of Sunnyvale, California; the ARM instruction set from ARM Holdings of Sunnyvale, California (with optional additional extensions such as NEON)), including the instruction(s) described herein. In one embodiment, core 1690 includes logic to support packed data instruction set extensions (e.g., AVX1, AVX2), allowing operations used by many multimedia applications to be performed using packed data.

[0203] It should be understood that a core may support multithreading (executing two or more parallel sets of operations or threads), and may do so in a variety of ways, including time-sliced ​​multithreading, simultaneous multithreading (where a single physical core provides a logical core for each of the threads that the physical core is simultaneously multithreading), or a combination thereof (e.g., time-sliced ​​fetching and decoding followed by simultaneous multithreading, such as in Hyper-Threading Technology).

[0204] Although register renaming is described in the context of out-of-order execution, it should be understood that register renaming can be used in an in-order architecture. Although the illustrated embodiment of the processor also includes separate instruction and data cache units 1634 / 1674 and a shared L2 cache unit 1676, alternative embodiments may have a single internal cache for both instructions and data, such as, for example, a level 1 (L1) internal cache, or multiple levels of internal caches. In some embodiments, the system may include a combination of internal caches and external caches external to the core and / or processor. Alternatively, all caches may be external to the core and / or processor.

[0205] Specific Exemplary In-Order Core Architectures

[0206] Fig.17A -B illustrates a block diagram of a more specific exemplary in-order core architecture, which would be one of several logic blocks in a chip (including other cores of the same type and / or different types). Depending on the application, the logic block communicates with some fixed function logic, memory I / O interfaces, and other necessary I / O logic through a high bandwidth interconnect network (e.g., a ring network).

[0207] Fig.17A 1704 and 1715. 1706 is a block diagram of a single processor core according to an embodiment of the present invention, the single processor core together with its connection to the on-die interconnect network 1702 and together with its local subset of the level 2 (L2) cache 1704; in one embodiment, the instruction decoder 1700 supports the x86 instruction set with a packed data instruction set extension. The L1 cache 1706 allows low latency access to cache memory into the scalar and vector units. Although in one embodiment (to simplify the design), the scalar unit 1708 and the vector unit 1710 use separate register sets (respectively, scalar registers 1712 and vector registers 1714), and data passed between them is written to memory and then read back from the level 1 (L1) cache 1706, alternative embodiments of the present invention may use different approaches (e.g., using a single register set or including a communication path that allows data to be passed between two register files without being written and read back).

[0208] The local subset of L2 cache 1704 is part of the global L2 cache, which is divided into separate local subsets, one for each processor core. Each processor core has a direct access path to its own local subset of L2 cache 1704. The data read by the processor core is stored in its L2 cache subset 1704 and can be quickly accessed, which is parallel to other processor cores accessing their own local L2 cache subsets. The data written by the processor core is stored in its own L2 cache subset 1704 and is flushed from other subsets, if necessary. The ring network ensures consistency for shared data. The ring network is bidirectional to allow agents, such as processor cores, L2 caches and other logic blocks to communicate with each other in the chip. Each ring data path is 1012 bits wide in each direction.

[0209] Fig. 17B According to an embodiment of the present invention Fig.17A An expanded view of a portion of a processor core in FIG. Fig. 17B The L1 data cache 1706A portion of the L1 cache 1704 is included, as well as more details about the vector unit 1710 and vector registers 1714. In particular, the vector unit 1710 is a 16-wide vector processing unit (VPU) (see 16-wide ALU 1728) that executes one or more of integer, single-precision floating, and double-precision floating instructions. The VPU supports swizzling register inputs using a swizzle unit 1720, numerical conversion using numerical conversion units 1722A-B, and copying on memory inputs using a copy unit 1724. Write mask registers 1726 allow vector writes that predict the result.

[0210] Fig.18 is a block diagram of a processor 1800 that may have more than one core, may have an integrated memory controller, and may have integrated graphics according to an embodiment of the present invention. Fig.18 The solid line box in the figure illustrates a processor 1800 having a single core 1802A, a system agent 1810, a set of one or more bus controller units 1816, while the dashed line box optionally added illustrates an alternative processor 1800 having multiple cores 1802A-N, a set of one or more integrated memory controller units 1814 in the system agent unit 1810, and dedicated logic 1808.

[0211] Thus, different implementations of processor 1800 may include: 1) a CPU, where the dedicated logic 1808 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores), and cores 1802A-N are one or more general purpose cores (e.g., general purpose in-order cores, general purpose out-of-order cores, a combination of the two); 2) a coprocessor, where cores 1802A-N are a large number of dedicated cores primarily intended for graphics and / or science (throughput); and 3) a coprocessor, where cores 1802A-N are a large number of general purpose in-order cores. Thus, processor 1800 may be a general purpose processor, a coprocessor, or a special purpose processor, such as, for example, a network or communication processor, a compression engine, a graphics processor, a GPGPU (general purpose graphics processing unit), a high throughput many integrated cores (MIC), a coprocessor (including 30 or more cores), an embedded processor, and the like. The processor may be implemented on one or more chips. Processor 1800 may be part of and / or may be implemented on one or more substrates, the substrate using a plurality of process technologies, such as, for example, any of BiCMOS, CMOS, or NMOS.

[0212] The memory hierarchy includes one or more levels of cache within the core, a set or one or more shared cache units 1806, and external memory (not shown) coupled to the set of integrated memory controller units 1814. The set of shared cache units 1806 may include one or more mid-level caches, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, a last level cache (LLC), and / or combinations thereof. While in one embodiment, a ring-based interconnect unit 1812 interconnects the integrated graphics logic 1808 (the integrated graphics logic 1808 is an example of dedicated logic and is also referred to herein as dedicated logic), the set of shared cache units 1806, and the system agent unit 1810 / (multiple) integrated memory controller units 1814, alternative embodiments may use any number of well-known techniques for interconnecting such units. In one embodiment, coherency is maintained between the one or more cache units 1806 and the cores 1802-AN.

[0213] In some embodiments, one or more of the cores 1802A-N can be multithreaded. The system agent 1810 includes those components that coordinate and operate the cores 1802A-N. The system agent unit 1810 may include, for example, a power control unit (PCU) and a display unit. The PCU may be or include logic and components required to regulate the power state of the cores 1802A-N and the integrated graphics logic 1808. The display unit is used to drive one or more externally connected displays.

[0214] The cores 1802A-N may be homogeneous or heterogeneous with respect to the architectural instruction set; that is, two or more of the cores 1802A-N may be capable of executing the same instruction set, while others may be capable of executing only a subset of that instruction set or a different instruction set.

[0215] Exemplary Computer Architecture

[0216] Figure 19-22 is a block diagram of an exemplary computer architecture. Other system designs and configurations known in the art for laptops, desktop computers, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, mobile phones, portable media players, handheld devices, and various other electronic devices are also suitable. Generally, a variety of systems or electronic devices that can incorporate processors and / or other execution logic as disclosed herein are generally suitable.

[0217] Reference now Fig.19 , a block diagram of a system 1900 according to one embodiment of the present invention is shown. The system 1900 may include one or more processors 1910, 1915, which are coupled to a controller hub 1920. In one embodiment, the controller hub 1920 includes a graphics memory controller hub (GMCH) 1990 and an input / output hub (IOH) 1950 (which may be on separate chips); the GMCH 1990 includes memory and a graphics controller to which a memory 1940 and a coprocessor 1945 are coupled; the IOH 1950 couples an input / output (I / O) device 1960 to the GMCH 1990. Alternatively, one or both of the memory and graphics controller are integrated within the processor (as described herein), the memory 1940 and the coprocessor 1945 are directly coupled to the processor 1910, and the controller hub 1920 in a single chip with the IOH 1950.

[0218] exist Fig.19 The optional nature of the additional processor 1915 is indicated by a broken line. Each processor 1910, 1915 may include one or more of the processing cores described herein, and may be a version of processor 1800.

[0219] The memory 1940 may be, for example, a dynamic random access memory (DRAM), a phase change memory (PCM), or a combination of the two. For at least one embodiment, the controller hub 1920 communicates with the processor(s) 1910, 1915 via a multi-drop bus, such as a front side bus (FSB), a point-to-point interface, such as a Quick Path Interconnect (QPI) or similar connection 1995.

[0220] In one embodiment, coprocessor 1945 is a special purpose processor such as, for example, a high throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, etc. In one embodiment, controller hub 1920 may include an integrated graphics accelerator.

[0221] There may be various differences between the physical resources 1910 , 1915 in terms of a range of metrics of merit including architectural, microarchitectural, thermal, power consumption characteristics, and the like.

[0222] In one embodiment, the processor 1910 executes instructions that control a general type of data processing operation. Embedded within the instructions may be coprocessor instructions. The processor 1910 recognizes these coprocessor instructions as being of a type that should be executed by the attached coprocessor 1945. Therefore, the processor 1910 issues these coprocessor instructions (or control signals representing coprocessor instructions) to the coprocessor 1945 on a coprocessor bus or other interconnect. The coprocessor(s) 1945 accept and execute the received coprocessor instructions.

[0223] Reference now Fig. 20 , which is a block diagram of a first more specific exemplary system 2000 according to an embodiment of the present invention. Fig. 20 As shown in FIG. 2 , multiprocessor system 2000 is a point-to-point interconnect system and includes a first processor 2070 and a second processor 2080 coupled via a point-to-point interconnect 2050. Each of processors 2070 and 2080 may be a version of processor 1800. In one embodiment of the present invention, processors 2070 and 2080 are processors 1910 and 1915, respectively, and coprocessor 2038 is coprocessor 1945. In another embodiment, processors 2070 and 2080 are processor 1910 and coprocessor 1945, respectively.

[0224] Processors 2070 and 2080 are shown, which include integrated memory controller (IMC) units 2072 and 2082, respectively. Processor 2070 also includes point-to-point (PP) interfaces 2076 and 2078 as part of its bus controller unit; similarly, the second processor 2080 includes PP interfaces 2086 and 2088. Processors 2070, 2080 can exchange information via point-to-point (PP) interface 2050 by using PP interface circuits 2078, 2088. Fig. 20 As shown in FIG. 2 , IMCs 2072 and 2082 couple the processors to respective memories, namely memory 2032 and memory 2034 , which may be portions of main memory locally attached to the respective processors.

[0225] The processors 2070, 2080 may each exchange information with the chipset 2090 via separate PP interfaces 2052, 2054 using point-to-point interface circuits 2076, 2094, 2086, 2098. The chipset 2090 may optionally exchange information with the coprocessor 2038 via a high performance interface 2092. In one embodiment, the coprocessor 2038 is a special purpose processor such as, for example, a high throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, and the like.

[0226] A shared cache (not shown) may be included in either processor or external to both processors, also connected to the processors via the PP interconnect so that if the processors are placed in a low power mode, local cache information of either or both processors may be stored in the shared cache.

[0227] Chipset 2090 may be coupled to first bus 2016 via interface 2096. In one embodiment, first bus 2016 may be a Peripheral Component Interconnect (PCI) bus, or a bus such as a PCI Express bus or another third generation I / O interconnect bus, although the scope of the invention is not so limited.

[0228] like Fig. 20As shown in , various I / O devices 2014 can be coupled to the first bus 2016 along with a bus bridge 2018, which couples the first bus 2016 to a second bus 2020. In one embodiment, one or more additional processors 2015, such as coprocessors, high throughput MIC processors, GPGPUs, accelerators (such as, for example, graphics accelerators or digital signal processing (DSP) units), field programmable gate arrays, or any other processors are coupled to the first bus 2016. In one embodiment, the second bus 2020 can be a low pin count (LPC) bus. In one embodiment, various devices can be coupled to the second bus 2020, including, for example, a keyboard and / or mouse 2022, communication devices 2027, and a storage unit 2028, such as a disk drive or other mass storage device, which may include instructions / code and data 2030. In addition, an audio I / O 2024 can be coupled to the second bus 2020. Note that other architectures are possible. For example, instead of Fig. 20 Instead of a point-to-point architecture, the system can implement a multi-point branch bus or other such architecture.

[0229] Reference now Fig.21 , shown is a block diagram of a second more specific exemplary system 2100 according to an embodiment of the present invention. Fig. 20 and 21 Like elements have like reference numerals, and Fig. 20 Some aspects of Fig.21 Omitted to avoid Fig.21 Other aspects are vague.

[0230] Fig.21 It is illustrated that processors 2070, 2080 may include integrated memory and I / O control logic ("CL") 2072 and 2082, respectively. Thus, CL 2072, 2082 includes an integrated memory controller unit and includes I / O control logic. Fig.21 It is illustrated that not only are memories 2032, 2034 coupled to the CL 2072, 2082, but also that I / O devices 2114 are coupled to the control logic 2072, 2082. Legacy I / O devices 2115 are coupled to the chipset 2090.

[0231] Reference now Fig. 22 , shown is a block diagram of a SoC 2200 according to an embodiment of the present invention. Fig.18 Similar elements in FIG. 1 have the same reference numerals. Also, dashed boxes are optional features on more advanced SoCs. Fig. 22, the interconnect unit(s) 2202 are coupled to: an application processor 2210, which includes a set of one or more cores 1802A-N, the cores including cache units 1804A-N, and (multiple) shared cache units 1806; a system agent unit 1810; (multiple) bus controller units 1816; (multiple) integrated memory controller units 1814; a set of one or more coprocessors 2220, which may include integrated graphics logic, graphics processors, audio processors, and video processors; static random access memory (SRAM) units 2230; direct memory access (DMA) units 2232; and a display unit 2240 for coupling to one or more external displays. In one embodiment, the coprocessor(s) 2220 include special purpose processors such as, for example, network or communication processors, compression engines, GPGPUs, high throughput MIC processors, embedded processors, and the like.

[0232] Embodiments of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of such implementation approaches. Embodiments of the present invention may be implemented as a computer program or program code executed on a programmable system, the programmable system comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0233] Program code, such as Fig. 20 The code 2030 illustrated in the figure can be applied to the input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For this application, a processing system includes any system having a processor, such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0234] Program code can be implemented with high-level procedural or object-oriented programming languages ​​to communicate with the processing system. If desired, program code can also be implemented with assembly or machine language. In fact, the mechanisms described herein are not limited to any specific programming language in scope. In any case, the language can be a compiled or interpreted language.

[0235] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium, which represent various logic within a processor, which when read by a machine causes the machine to be made into logic for performing the techniques described herein. Such representations, known as "IP cores," may be stored on a tangible machine-readable medium and supplied to various customers or manufacturing facilities to be loaded into a manufacturing machine that actually constructs the logic or processor.

[0236] Such machine-readable storage media may include, without limitation, a non-transitory, tangible arrangement of an article made or formed by a machine or device, including storage media such as a hard disk, any other type of disk, including floppy disks, optical disks, compact disk read-only memory (CD-ROM), compact disk rewritable (CD-RW), and magneto-optical disks, semiconductor devices, such as read-only memory (ROM), random access memory (RAM), such as dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM), phase change memory (PCM), magnetic or optical cards, or any other type of medium suitable for storing electronic instructions.

[0237] Therefore, embodiments of the present invention also include non-transitory, tangible machine-readable media containing instructions or containing design data, such as hardware description language (HDL), which defines the structures, circuits, devices, processors and / or system features described herein. Such embodiments may also be referred to as program products.

[0238] Emulation (including binary conversion, code deformation, etc.)

[0239] In some cases, the instruction converter can be used to convert instructions from a source instruction set to a target instruction set. For example, the instruction converter can convert (e.g., by using static binary translation, dynamic binary translation, including dynamic compilation), deform, simulate or otherwise convert instructions into one or more other instructions to be processed by the core. The instruction converter can be implemented in software, hardware, firmware or a combination thereof. The instruction converter can be on the processor, away from the processor, or partially on the processor and partially away from the processor.

[0240] Fig.23 1 is a block diagram that compares the use of a software instruction converter to convert binary instructions in a source instruction set into binary instructions in a target instruction set according to an embodiment of the present invention. In the illustrated embodiment, the instruction converter is a software instruction converter, although alternatively, the instruction converter can be implemented in software, firmware, hardware, or various combinations thereof. Fig.23It is shown that a program in a high-level language 2302 can be compiled using an x86 compiler 2304 to generate x86 binary code 2306, which can be natively executed by a processor having at least one x86 instruction set core 2316. A processor having at least one x86 instruction set core 2316 represents any processor that can perform substantially the same functions as an Intel processor having at least one x86 instruction set core by compatibly executing or otherwise processing (1) a substantial portion of the instruction set of an Intel x86 instruction set core, or (2) an object code version of an application or other software targeted to run on an Intel processor having at least one x86 instruction set core, so as to achieve substantially the same results as an Intel processor having at least one x86 instruction set core. x86 compiler 2304 represents a compiler operable to generate x86 binary code 2306 (e.g., object code) that can be executed on a processor having at least one x86 instruction set core 2316 with or without additional link processing. Similarly, Fig.23 It is shown that a program in a high-level language 2302 can be compiled using an alternative instruction set compiler 2308 to generate an alternative instruction set binary code 2310 that can be natively executed by a processor that does not have at least one x86 instruction set core 2314 (e.g., a processor having a core that executes the MIPS instruction set of MIPS Technologies of Sunnyvale, California and / or executes the ARM instruction set of ARM Holdings of Sunnyvale, California). An instruction converter 2312 is used to convert the x86 binary code 2306 into code that can be natively executed by a processor that does not have an x86 instruction set core 2314. Such converted code is unlikely to be the same as the alternative instruction set binary code 2310 because an instruction converter that can do so is difficult to make; however, the converted code will implement general operations and be composed of instructions from the alternative instruction set. Thus, instruction converter 2312 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device that does not have an x86 instruction set processor or core to execute x86 binary code 2306 through simulation, emulation, or any other process.

Claims

1. A device comprising: decoder logic configured to decode a single instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein packed data elements of the first packed data source operand and the second packed data source operand have a first size and packed data elements of the packed data source / destination operand have a second size that is larger than the first size; a register file having a plurality of packed data registers for storing one or more of the packed data source / destination operands, the first packed data source operand, and the second packed data source operand; and execution logic coupled to the decoder logic and the register file, wherein, in response to a decoded single instruction, the execution logic is configured, according to the opcode of the single instruction, for each packed data element position of the packed data source / destination operand, to: sign extending a plurality of packed data bytes from corresponding packed data element positions of the first packed data source operand; zero-extending a plurality of packed data bytes at corresponding packed data element positions from the second packed data source operand; multiplying each of the sign-extended plurality of packed data bytes from the first packed data source operand with a corresponding one of the zero-extended plurality of packed data bytes from the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result; and The addition result is stored in the corresponding packed data element position of the packed data source / destination operand.

2. The device according to claim 1, wherein: The execution logic is configured to suppress memory failures.

3. The device according to claim 1, wherein: When the single instruction further includes another field for a write mask, the execution logic is configured to perform a merge operation.

4. The device according to any one of claims 1 to 3, wherein: The execution logic is configured to sign extend a plurality of packed data bytes from the first packed data source operand, the plurality of packed data bytes from the first packed data source operand including a signed byte.

5. The device according to any one of claims 1 to 3, wherein: The execution logic is configured to zero-extend a plurality of packed data bytes from the second packed data source operand, the plurality of packed data bytes from the second packed data source operand comprising unsigned bytes.

6. The device according to any one of claims 1 to 3, wherein: When the packed data source / destination operand has a width of 128 bits, the execution logic is configured to perform four iterations of the multiplication, the addition, and the storage.

7. The device according to any one of claims 1 to 3, wherein: When the packed data source / destination operand has a width of 256 bits, the execution logic is configured to perform 8 iterations of the multiplication, the addition, and the storage.

8. A method comprising: decoding, in a decoder of the processor, a single instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein packed data elements of the first packed data source operand and the second packed data source operand have a first size and packed data elements of the packed data source / destination operand have a second size that is larger than the first size; and In execution logic coupled to the decoder, performing the following operations for each packed data element position of the packed data source / destination operand according to the opcode of the single instruction: sign extending a plurality of packed data bytes from corresponding packed data element positions of the first packed data source operand; zero-extending a plurality of packed data bytes at corresponding packed data element positions from the second packed data source operand; multiplying each of the sign-extended plurality of packed data bytes from the first packed data source operand with a corresponding one of the zero-extended plurality of packed data bytes from the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result; and The addition result is stored in the corresponding packed data element position of the packed data source / destination operand.

9. The method of claim 8, wherein: The executing further includes suppressing a memory fault.

10. The method of claim 8, wherein: The executing further includes performing a merge operation when the single instruction further includes another field for a write mask.

11. The method according to any one of claims 8 to 10, wherein: The executing further includes sign extending a plurality of packed data bytes from the first packed data source operand, the plurality of packed data bytes from the first packed data source operand including a signed byte.

12. The method according to any one of claims 8 to 10, wherein: The executing further includes zero-extending a plurality of packed data bytes from the second packed data source operand, the plurality of packed data bytes from the second packed data source operand including unsigned bytes.

13. The method according to any one of claims 8 to 10, wherein: The executing further includes performing 4 iterations of the multiplying, the adding, and the storing when the packed data source / destination operand has a width of 128 bits.

14. The method according to any one of claims 8 to 10, wherein: The executing further includes performing 8 iterations of the multiplying, the adding, and the storing when the packed data source / destination operand has a width of 256 bits.

15. A non-transitory machine-readable medium comprising instructions that, when executed by a processor, cause the processor to: The instruction is decoded in a decoder of the processor, the instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein: The packed data elements of the first packed data source operand and the second packed data source operand have a first size, and the packed data elements of the packed data source / destination operand have a second size that is larger than the first size; as well as In execution logic of the processor coupled to the decoder, performing the following operations for each packed data element position of the packed data source / destination operand according to the opcode of the instruction: sign extending a plurality of packed data bytes from corresponding packed data element positions of the first packed data source operand; zero-extending a plurality of packed data bytes at corresponding packed data element positions from the second packed data source operand; multiplying each of the sign-extended plurality of packed data bytes from the first packed data source operand with a corresponding one of the zero-extended plurality of packed data bytes from the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result; as well as The addition result is stored in the corresponding packed data element position of the packed data source / destination operand.

16. The non-transitory machine-readable medium of claim 15, wherein: The instructions, when executed by the processor, cause the processor to suppress memory faults.

17. The non-transitory machine-readable medium of claim 15, wherein: The instructions, when executed by the processor, cause the processor to: perform a merge operation when the instructions further include another field for a write mask.

18. The non-transitory machine-readable medium of any one of claims 15 to 17, wherein: The instructions, when executed by the processor, cause the processor to sign extend a plurality of packed data bytes from the first packed data source operand, the plurality of packed data bytes from the first packed data source operand including a signed byte.

19. The non-transitory machine-readable medium of any one of claims 15 to 17, wherein: The instructions, when executed by the processor, cause the processor to zero-extend a plurality of packed data bytes from the second packed data source operand, the plurality of packed data bytes from the second packed data source operand comprising unsigned bytes.

20. The non-transitory machine-readable medium of any one of claims 15 to 17, wherein: The instructions, when executed by the processor, cause the processor to: perform 4 iterations of the multiplication, the addition, and the storage when the packed data source / destination operand has a width of 128 bits.

21. The non-transitory machine-readable medium of any one of claims 15 to 17, wherein: The instructions, when executed by the processor, cause the processor to: perform 8 iterations of the multiplication, the addition, and the storage when the packed data source / destination operand has a width of 256 bits.

22. A system comprising: A processor, the processor comprising: a decoder for decoding a single instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein packed data elements of the first packed data source operand and the second packed data source operand have a first size and packed data elements of the packed data source / destination operand have a second size that is larger than the first size; a register file having a plurality of packed data registers for storing one or more of the packed data source / destination operands, the first packed data source operand, and the second packed data source operand; and and execution logic coupled to the decoder and the register file, wherein, in response to a decoded single instruction, the execution logic is to, for each packed data element position of the packed data source / destination operand according to the opcode of the single instruction: sign extending a plurality of packed data bytes from corresponding packed data element positions of the first packed data source operand; zero-extending a plurality of packed data bytes at corresponding packed data element positions from the second packed data source operand; multiplying each of the sign-extended plurality of packed data bytes from the first packed data source operand with a corresponding one of the zero-extended plurality of packed data bytes from the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result; and storing the addition result in the corresponding packed data element position of the packed data source / destination operand; and A dynamic random access memory is coupled to the processor.

23. The system of claim 22, wherein: The execution logic is used to suppress memory failures.

24. The system of claim 22, wherein: When the single instruction further includes another field for a write mask, the execution logic is configured to perform a merge operation.

25. A system as claimed in any one of claims 22 to 24, wherein: The execution logic is to sign extend a plurality of packed data bytes from the first packed data source operand, the plurality of packed data bytes from the first packed data source operand including a signed byte.

26. The system of any one of claims 22 to 24, wherein: The execution logic is to zero-extend a plurality of packed data bytes from the second packed data source operand, the plurality of packed data bytes from the second packed data source operand comprising unsigned bytes.

27. The system of any one of claims 22 to 24, wherein: When the packed data source / destination operand has a width of 128 bits, the execution logic is to perform 4 iterations of the multiplication, the addition, and the storage.

28. The system of any one of claims 22 to 24, wherein: When the packed data source / destination operand has a width of 256 bits, the execution logic is to perform 8 iterations of the multiplication, the addition, and the storage.

29. An apparatus comprising: decoder logic configured to decode a single instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein packed data elements of the first packed data source operand and the second packed data source operand have a first size and packed data elements of the packed data source / destination operand have a second size that is larger than the first size; a register file having a plurality of packed data registers for storing one or more of the packed data source / destination operands, the first packed data source operand, and the second packed data source operand; and execution logic coupled to the decoder logic and the register file, wherein, in response to a decoded single instruction, the execution logic is configured, according to the opcode of the single instruction, for each packed data element position of the packed data source / destination operand, to: sign extending a plurality of packed data words at corresponding packed data element positions from the first packed data source operand; sign extending a plurality of packed data words at corresponding packed data element positions from the second packed data source operand; multiplying each of a plurality of sign-extended packed data words from a corresponding packed data element position of the first packed data source operand with a corresponding one of a plurality of sign-extended packed data words from a corresponding packed data element position of the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result; and The addition result is stored in the corresponding packed data element position of the packed data source / destination operand.

30. The apparatus of claim 29, wherein: The execution logic is configured to suppress memory failures.

31. The apparatus of claim 29, wherein: When the single instruction further includes another field for a write mask, the execution logic is configured to perform a merge operation.

32. The device of any one of claims 29 to 31, wherein: The execution logic is configured to sign extend a plurality of packed data words from the first packed data source operand, the plurality of packed data words from the first packed data source operand comprising signed words.

33. The device of any one of claims 29 to 31, wherein: The execution logic is configured to generate the addition result including a doubleword result.

34. The device of any one of claims 29 to 31, wherein: When the packed data source / destination operand has a width of 128 bits, the execution logic is configured to perform four iterations of the multiplication, the addition, and the storage.

35. The device of any one of claims 29 to 31, wherein: When the packed data source / destination operand has a width of 256 bits, the execution logic is configured to perform 8 iterations of the multiplication, the addition, and the storage.

36. A method comprising: decoding, in a decoder of the processor, a single instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein packed data elements of the first packed data source operand and the second packed data source operand have a first size and packed data elements of the packed data source / destination operand have a second size that is larger than the first size; and In execution logic of the processor coupled to the decoder, performing the following operations for each packed data element position of the packed data source / destination operand according to the opcode of the single instruction: sign extending a plurality of packed data words at corresponding packed data element positions from the first packed data source operand; sign extending a plurality of packed data words at corresponding packed data element positions from the second packed data source operand; multiplying each of a plurality of sign-extended packed data words from a corresponding packed data element position of the first packed data source operand with a corresponding one of a plurality of sign-extended packed data words from a corresponding packed data element position of the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result; and The addition result is stored in the corresponding packed data element position of the packed data source / destination operand.

37. The method of claim 36, wherein: The executing further includes suppressing a memory fault.

38. The method of claim 36, wherein: The executing further includes performing a merge operation when the single instruction further includes another field for a write mask.

39. The method of any one of claims 36 to 38, wherein: The executing further includes sign extending a plurality of packed data words from the first packed data source operand, the plurality of packed data words from the first packed data source operand including signed words.

40. The method of any one of claims 36 to 38, wherein: The executing further includes generating the addition result comprising a doubleword.

41. The method of any one of claims 36 to 38, wherein: The executing further includes performing 4 iterations of the multiplying, the adding, and the storing when the packed data source / destination operand has a width of 128 bits.

42. The method of any one of claims 36 to 38, wherein: The executing further includes performing 8 iterations of the multiplying, the adding, and the storing when the packed data source / destination operand has a width of 256 bits.

43. A non-transitory machine-readable medium comprising instructions that, when executed by a processor, cause the processor to: The instruction is decoded in a decoder of the processor, the instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein: The packed data elements of the first packed data source operand and the second packed data source operand have a first size, and the packed data elements of the packed data source / destination operand have a second size that is larger than the first size; as well as In execution logic of the processor coupled to the decoder, performing the following operations for each packed data element position of the packed data source / destination operand according to the opcode of the instruction: sign extending a plurality of packed data words at corresponding packed data element positions from the first packed data source operand; sign extending a plurality of packed data words at corresponding packed data element positions from the second packed data source operand; multiplying each of a plurality of sign-extended packed data words from a corresponding packed data element position of the first packed data source operand with a corresponding one of a plurality of sign-extended packed data words from a corresponding packed data element position of the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result; as well as The addition result is stored in the corresponding packed data element position of the packed data source / destination operand.

44. The non-transitory machine-readable medium of claim 43, wherein: The instructions, when executed by the processor, cause the processor to suppress memory faults.

45. The non-transitory machine-readable medium of claim 43, wherein: The instructions, when executed by the processor, cause the processor to: perform a merge operation when the instructions further include another field for a write mask.

46. ​​The non-transitory machine-readable medium of any one of claims 43 to 45, wherein: The instructions, when executed by the processor, cause the processor to sign-extend a plurality of packed data words from the first packed data source operand, the plurality of packed data words from the first packed data source operand including signed words.

47. The non-transitory machine-readable medium of any one of claims 43 to 45, wherein: The instructions, when executed by the processor, cause the processor to: perform 4 iterations of the multiplication, the addition, and the storage when the packed data source / destination operand has a width of 128 bits.

48. The non-transitory machine-readable medium of any one of claims 43 to 45, wherein: The instructions, when executed by the processor, cause the processor to: perform 8 iterations of the multiplication, the addition, and the storage when the packed data source / destination operand has a width of 256 bits.

49. A system comprising: A processor, the processor comprising: a decoder for decoding a single instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein packed data elements of the first packed data source operand and the second packed data source operand have a first size and packed data elements of the packed data source / destination operand have a second size that is larger than the first size; a register file having a plurality of packed data registers for storing one or more of the packed data source / destination operands, the first packed data source operand, and the second packed data source operand; and and execution logic coupled to the decoder and the register file, wherein, in response to a decoded single instruction, the execution logic is to, for each packed data element position of the packed data source / destination operand according to the opcode of the single instruction: sign extending a plurality of packed data words at corresponding packed data element positions from the first packed data source operand; sign extending a plurality of packed data words at corresponding packed data element positions from the second packed data source operand; multiplying each of a plurality of sign-extended packed data words from a corresponding packed data element position of the first packed data source operand with a corresponding one of a plurality of sign-extended packed data words from a corresponding packed data element position of the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result; and storing the addition result in the corresponding packed data element position of the packed data source / destination operand; and A dynamic random access memory is coupled to the processor.

50. The system of claim 49, wherein: The execution logic is used to suppress memory failures.

51. The system of claim 49, wherein: When the single instruction further includes another field for a write mask, the execution logic is configured to perform a merge operation.

52. A system as claimed in any one of claims 49 to 51, wherein: The execution logic is to sign extend a plurality of packed data words from the first packed data source operand, the plurality of packed data words from the first packed data source operand including signed words.

53. A system as claimed in any one of claims 49 to 51, wherein: When the packed data source / destination operand has a width of 128 bits, the execution logic is to perform 4 iterations of the multiplication, the addition, and the storage.

54. A system as claimed in any one of claims 49 to 51, wherein: When the packed data source / destination operand has a width of 256 bits, the execution logic is to perform 8 iterations of the multiplication, the addition, and the storage.

55. An apparatus comprising: decoder logic configured to decode a single instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein packed data elements of the first packed data source operand and the second packed data source operand have a first size and packed data elements of the packed data source / destination operand have a second size; a register file having a plurality of packed data registers for storing one or more of the packed data source / destination operands, the first packed data source operand, and the second packed data source operand; and execution logic coupled to the decoder logic and the register file, wherein, in response to a decoded single instruction, the execution logic is configured, according to the opcode of the single instruction, for each packed data element position of the packed data source / destination operand, to: sign-extending a plurality of packed signed data bytes from corresponding packed data element positions of the first packed data source operand; zero-extending a plurality of packed unsigned data bytes at corresponding packed data element positions from the second packed data source operand; multiplying each of the sign-extended plurality of packed signed data bytes from the first packed data source operand with a corresponding one of the zero-extended plurality of packed unsigned data bytes from the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result, and saturating the addition result if a width of the addition result exceeds a width of the second size to produce a saturated addition result; and The addition result or the saturated addition result is stored in the corresponding packed data element position of the packed data source / destination operand.

56. The apparatus of claim 55, wherein: The execution logic is configured to suppress memory failures.

57. The apparatus of claim 55, wherein: When the single instruction further includes another field for a write mask, the execution logic is configured to perform a merge operation.

58. The apparatus of any one of claims 55 to 57, wherein: When the packed data source / destination operand has a width of 128 bits, the execution logic is configured to perform four iterations of the multiplication, the addition, and the storage.

59. The device of any one of claims 55 to 57, wherein: When the packed data source / destination operand has a width of 256 bits, the execution logic is configured to perform 8 iterations of the multiplication, the addition, and the storage.

60. A method comprising: decoding, in a decoder of the processor, a single instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein packed data elements of the first packed data source operand and the second packed data source operand have a first size and packed data elements of the packed data source / destination operand have a second size; and In execution logic coupled to the decoder, performing the following operations for each packed data element position of the packed data source / destination operand according to the opcode of the single instruction: sign-extending a plurality of packed signed data bytes from corresponding packed data element positions of the first packed data source operand; zero-extending a plurality of packed unsigned data bytes at corresponding packed data element positions from the second packed data source operand; multiplying each of the sign-extended plurality of packed signed data bytes from the first packed data source operand with a corresponding one of the zero-extended plurality of packed unsigned data bytes from the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result, and saturating the addition result if a width of the addition result exceeds a width of the second size to produce a saturated addition result; and The addition result or the saturated addition result is stored in the corresponding packed data element position of the packed data source / destination operand.

61. The method of claim 60, wherein: The executing further includes suppressing a memory fault.

62. The method of claim 60, wherein: The executing further includes performing a merge operation when the single instruction further includes another field for a write mask.

63. The method of any one of claims 60 to 62, wherein: The executing further includes performing 4 iterations of the multiplying, the adding, and the storing when the packed data source / destination operand has a width of 128 bits.

64. The method of any one of claims 60 to 62, wherein: The executing further includes performing 8 iterations of the multiplying, the adding, and the storing when the packed data source / destination operand has a width of 256 bits.

65. A non-transitory machine-readable medium comprising instructions that, when executed by a processor, cause the processor to: The instruction is decoded in a decoder of the processor, the instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein: The packed data elements of the first packed data source operand and the second packed data source operand have a first size, and the packed data elements of the packed data source / destination operand have a second size; and In execution logic of the processor coupled to the decoder, performing the following operations for each packed data element position of the packed data source / destination operand according to the opcode of the instruction: sign-extending a plurality of packed signed data bytes from corresponding packed data element positions of the first packed data source operand; zero-extending a plurality of packed unsigned data bytes at corresponding packed data element positions from the second packed data source operand; multiplying each of the sign-extended plurality of packed signed data bytes from the first packed data source operand with a corresponding one of the zero-extended plurality of packed unsigned data bytes from the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result, and saturating the addition result if a width of the addition result exceeds a width of the second size to produce a saturated addition result; as well as The addition result or the saturated addition result is stored in the corresponding packed data element position of the packed data source / destination operand.

66. The non-transitory machine-readable medium of claim 65, wherein: The instructions, when executed by the processor, cause the processor to suppress memory faults.

67. The non-transitory machine-readable medium of claim 65, wherein: The instructions, when executed by the processor, cause the processor to: perform a merge operation when the instructions further include another field for a write mask.

68. The non-transitory machine-readable medium of any one of claims 65 to 67, wherein: The instructions, when executed by the processor, cause the processor to: perform 4 iterations of the multiplication, the addition, and the storage when the packed data source / destination operand has a width of 128 bits.

69. The non-transitory machine-readable medium of any one of claims 65 to 67, wherein: The instructions, when executed by the processor, cause the processor to: perform 8 iterations of the multiplication, the addition, and the storage when the packed data source / destination operand has a width of 256 bits.

70. A system comprising: A processor, the processor comprising: a decoder for decoding a single instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein packed data elements of the first packed data source operand and the second packed data source operand have a first size and packed data elements of the packed data source / destination operand have a second size; a register file having a plurality of packed data registers for storing one or more of the packed data source / destination operands, the first packed data source operand, and the second packed data source operand; and and execution logic coupled to the decoder and the register file, wherein, in response to a decoded single instruction, the execution logic is to, for each packed data element position of the packed data source / destination operand according to the opcode of the single instruction: sign-extending a plurality of packed signed data bytes from corresponding packed data element positions of the first packed data source operand; zero-extending a plurality of packed unsigned data bytes at corresponding packed data element positions from the second packed data source operand; multiplying each of the sign-extended plurality of packed signed data bytes from the first packed data source operand with a corresponding one of the zero-extended plurality of packed unsigned data bytes from the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result, and saturating the addition result if a width of the addition result exceeds a width of the second size to produce a saturated addition result; and storing the addition result or the saturated addition result in the corresponding packed data element position of the packed data source / destination operand; and A dynamic random access memory is coupled to the processor.

71. The system of claim 70, wherein: The execution logic is used to suppress memory failures.

72. The system of claim 70, wherein: When the single instruction further includes another field for a write mask, the execution logic is configured to perform a merge operation.

73. The system of any one of claims 70 to 72, wherein: When the packed data source / destination operand has a width of 128 bits, the execution logic is to perform 4 iterations of the multiplication, the addition, and the storage.

74. The system of any one of claims 70 to 72, wherein: When the packed data source / destination operand has a width of 256 bits, the execution logic is to perform 8 iterations of the multiplication, the addition, and the storage.

75. An apparatus comprising: decoder logic configured to decode a single instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein packed data elements of the first packed data source operand and the second packed data source operand have a first size and packed data elements of the packed data source / destination operand have a second size; a register file having a plurality of packed data registers for storing one or more of the packed data source / destination operands, the first packed data source operand, and the second packed data source operand; and execution logic coupled to the decoder logic and the register file, wherein, in response to a decoded single instruction, the execution logic is configured, according to the opcode of the single instruction, for each packed data element position of the packed data source / destination operand, to: sign-extending a plurality of packed signed data words from corresponding packed data element positions of the first packed data source operand; sign-extending a plurality of packed signed data words from corresponding packed data element positions of the second packed data source operand; multiplying each of a plurality of sign-extended packed signed data words from a corresponding packed data element position of the first packed data source operand with a corresponding one of a plurality of sign-extended packed signed data words from a corresponding packed data element position of the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result, and saturating the addition result if a width of the addition result exceeds a width of the second size to produce a saturated addition result; and The addition result or the saturated addition result is stored in the corresponding packed data element position of the packed data source / destination operand.

76. The apparatus of claim 75, wherein: The execution logic is configured to suppress memory failures.

77. The apparatus of claim 75, wherein: When the single instruction further includes another field for a write mask, the execution logic is configured to perform a merge operation.

78. The device of any one of claims 75 to 77, wherein: When the packed data source / destination operand has a width of 128 bits, the execution logic is configured to perform four iterations of the multiplication, the addition, and the storage.

79. The device of any one of claims 75 to 77, wherein: When the packed data source / destination operand has a width of 256 bits, the execution logic is configured to perform 8 iterations of the multiplication, the addition, and the storage.

80. A method comprising: decoding, in a decoder of the processor, a single instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein packed data elements of the first packed data source operand and the second packed data source operand have a first size and packed data elements of the packed data source / destination operand have a second size; and In execution logic of the processor coupled to the decoder, performing the following operations for each packed data element position of the packed data source / destination operand according to the opcode of the single instruction: sign-extending a plurality of packed signed data words from corresponding packed data element positions of the first packed data source operand; sign-extending a plurality of packed signed data words from corresponding packed data element positions of the second packed data source operand; multiplying each of a plurality of sign-extended packed signed data words from a corresponding packed data element position of the first packed data source operand with a corresponding one of a plurality of sign-extended packed signed data words from a corresponding packed data element position of the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result, and saturating the addition result if a width of the addition result exceeds a width of the second size to produce a saturated addition result; and The addition result or the saturated addition result is stored in the corresponding packed data element position of the packed data source / destination operand.

81. The method of claim 80, wherein: The executing further includes suppressing a memory fault.

82. The method of claim 80, wherein: The executing further includes performing a merge operation when the single instruction further includes another field for a write mask.

83. The method of any one of claims 80 to 82, wherein: The executing further includes performing 4 iterations of the multiplying, the adding, and the storing when the packed data source / destination operand has a width of 128 bits.

84. The method of any one of claims 80 to 82, wherein: The executing further includes performing 8 iterations of the multiplying, the adding, and the storing when the packed data source / destination operand has a width of 256 bits.

85. A non-transitory machine-readable medium comprising instructions that, when executed by a processor, cause the processor to: The instruction is decoded in a decoder of the processor, the instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein: The packed data elements of the first packed data source operand and the second packed data source operand have a first size, and the packed data elements of the packed data source / destination operand have a second size; and In execution logic of the processor coupled to the decoder, performing the following operations for each packed data element position of the packed data source / destination operand according to the opcode of the instruction: sign-extending a plurality of packed signed data words from corresponding packed data element positions of the first packed data source operand; sign-extending a plurality of packed signed data words from corresponding packed data element positions of the second packed data source operand; multiplying each of a plurality of sign-extended packed signed data words from a corresponding packed data element position of the first packed data source operand with a corresponding one of a plurality of sign-extended packed signed data words from a corresponding packed data element position of the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result, and saturating the addition result if a width of the addition result exceeds a width of the second size to produce a saturated addition result; as well as The addition result or the saturated addition result is stored in the corresponding packed data element position of the packed data source / destination operand.

86. The non-transitory machine-readable medium of claim 85, wherein: The instructions, when executed by the processor, cause the processor to suppress memory faults.

87. The non-transitory machine-readable medium of claim 85, wherein: The instructions, when executed by the processor, cause the processor to: perform a merge operation when the instructions further include another field for a write mask.

88. The non-transitory machine-readable medium of any one of claims 85 to 87, wherein: The instructions, when executed by the processor, cause the processor to: perform 4 iterations of the multiplication, the addition, and the storage when the packed data source / destination operand has a width of 128 bits.

89. The non-transitory machine-readable medium of any one of claims 85 to 87, wherein: The instructions, when executed by the processor, cause the processor to: perform 8 iterations of the multiplication, the addition, and the storage when the packed data source / destination operand has a width of 256 bits.

90. A system comprising: A processor, the processor comprising: a decoder for decoding a single instruction having an opcode, a first field for representing a packed data source / destination operand, a second field for representing a first packed data source operand, and a third field for representing a second packed data source operand, wherein packed data elements of the first packed data source operand and the second packed data source operand have a first size and packed data elements of the packed data source / destination operand have a second size; a register file having a plurality of packed data registers for storing one or more of the packed data source / destination operands, the first packed data source operand, and the second packed data source operand; and and execution logic coupled to the decoder and the register file, wherein, in response to a decoded single instruction, the execution logic is to, for each packed data element position of the packed data source / destination operand according to the opcode of the single instruction: sign-extending a plurality of packed signed data words from corresponding packed data element positions of the first packed data source operand; sign-extending a plurality of packed signed data words from corresponding packed data element positions of the second packed data source operand; multiplying each of a plurality of sign-extended packed signed data words from a corresponding packed data element position of the first packed data source operand with a corresponding one of a plurality of sign-extended packed signed data words from a corresponding packed data element position of the second packed data source operand to produce a plurality of results; adding the plurality of results to packed data elements of the second size at corresponding packed data element positions of the packed data source / destination operand to produce an addition result, and saturating the addition result if a width of the addition result exceeds a width of the second size to produce a saturated addition result; and storing the addition result or the saturated addition result in the corresponding packed data element position of the packed data source / destination operand; and A dynamic random access memory is coupled to the processor.

91. The system of claim 90, wherein: The execution logic is used to suppress memory failures.

92. The system of claim 90, wherein: When the single instruction further includes another field for a write mask, the execution logic is configured to perform a merge operation.

93. The system of any one of claims 90 to 92, wherein: When the packed data source / destination operand has a width of 128 bits, the execution logic is to perform 4 iterations of the multiplication, the addition, and the storage.

94. The system of any one of claims 90 to 92, wherein: When the packed data source / destination operand has a width of 256 bits, the execution logic is to perform 8 iterations of the multiplication, the addition, and the storage.

Citation Information

Patent Citations

  • Multiple Data Element-To-Multiple Data Element Comparison Processors, Methods, Systems, and Instructions

    CN104049954A

  • Dot product processors, methods, systems, and instructions

    CN104137055A