System and method for executing instructions to convert into 16-bit floating point format

By introducing extraction, decoding and execution circuits into the processor, efficient conversion of single-precision vectors to 16-bit floating-point format is realized, solving the problem of insufficient performance and power efficiency in the prior art, and improving the performance of machine learning applications.

CN120540718APending Publication Date: 2025-08-26INTEL CORP
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510595068.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-11-09
Filing Date
2019-10-30
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently convert to a 16-bit floating point format when processing single-precision vectors, resulting in insufficient performance and power efficiency, especially in machine learning applications.

Method used

A processor is provided, including an extraction circuit, a decoding circuit and an execution circuit, capable of extracting and decoding instructions, converting single-precision elements into a 16-bit floating-point format through an opcode, and truncating and rounding when necessary, and storing them into a target vector.

Benefits of technology

Achieves the same quality as single-precision algorithms, but reduces memory usage and memory bandwidth requirements, and improves performance and power efficiency, especially in machine learning environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540718A_ABST
    Figure CN120540718A_ABST
Patent Text Reader

Abstract

The disclosure relates to systems and methods for executing instructions to convert into a 16-bit floating point format. In one embodiment, a processor includes fetch circuitry to fetch an instruction having a field to specify positions of an opcode and a first source vector including N single precision elements and a target vector including at least N 16-bit floating point elements, the opcode instructs execution circuitry to convert each element of the first source vector into a 16-bit floating point format and store each converted element into a corresponding position of the target vector, the conversion including truncation and rounding if necessary; the decoding circuit is used for decoding the instruction; and execution circuitry to respond to the instruction in accordance with the opcode.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Description of the case

[0002] This application is a divisional application of the invention patent application with application number 201911045764.1 and name “System and method for executing instructions to convert into 16-bit floating point format”. Technical Field

[0003] The field of the invention relates generally to computer processor architecture and, more particularly, to systems and methods for executing instructions to convert to a 16-bit floating point format. Background Art

[0004] An instruction set, or instruction set architecture (ISA), is the part of a computer architecture that is relevant to programming and may include native data types, instructions, register architecture, addressing modes, memory architecture, interrupt and exception handling, and external input and output (I / O). An instruction set includes one or more instruction formats. A given instruction format defines various fields (number of bits, position of bits) to specify the operation to be performed and the operand(s) on which the operation is to be performed, etc. A given instruction is expressed using a given instruction format and specifies the operation and operand. An instruction stream is a specific sequence of instructions, where each instruction in the sequence is an occurrence of a certain instruction in a certain instruction format.

[0005] Scientific, financial, general-purpose, auto-vectorized, RMS (recognition, mining, and synthesis) / vision, and multimedia applications (e.g., 2D / 3D graphics, image processing, video compression / decompression, speech recognition algorithms, and audio manipulation) often require the same operation to be performed on a large number of data items (referred to as "data parallelism"). Single Instruction Multiple Data (SIMD) refers to a class of instructions that causes a processor to perform the same operation on multiple data items. SIMD technology is particularly well-suited for processors that can logically divide the bits in a register into several fixed-size data elements, each representing a separate value. For example, the bits in a 512-bit register can be specified as source operands to be operated on as sixteen separate 32-bit single-precision floating-point data elements. As another example, the bits in a 256-bit register can be specified as source operands to be operated on as sixteen separate 16-bit floating-point packed data elements, eight separate 32-bit packed data elements (doubleword-sized data elements), or thirty-two separate 8-bit data elements (byte (B)-sized data elements). This type of data is called a packed data type or vector data type, and operands of this data type are called packed data operands or vector operands. In other words, a packed data item or vector refers to a sequence of packed data elements; and a packed data operand or vector operand is the source or destination operand of a SIMD instruction (also called a packed data instruction or vector instruction).

[0006] As an example, a class of SIMD instructions specifies the single vector operation to be performed on two source vector operands in a vertical manner, to generate the target vector operand with the same size, with the data elements of the same number and by the same data element order.The data element in the source vector operand is called the source data element, and the data element in the target vector operand is called the target or result data element. These source vector operands have the same size and comprise the data elements of same width, thereby they comprise the data elements of the same number. The source data element that is in the same bit position in the two source vector operands forms the pair (also referred to as corresponding data element; That is to say, the data element in the data element position 0 of each source operand is corresponding, and the data element in the data element position 1 of each source operand is corresponding, etc. and so on). The operation specified by this SIMD instruction is performed respectively on each pair of these pairs of source data elements to generate the result data element of matching number, thereby every pair of source data elements has corresponding result data element. Because the operations are vertical and because the result vector operands are the same size, have the same number of data elements, and the result data elements are stored in the same data element order as the source vector operands, the result data elements are in the same bit positions in the result vector operands as their corresponding pairs of source data elements in the source vector operands. In addition to this exemplary type of SIMD instruction, various other types of SIMD instructions exist.

[0007] Some applications that process vectors with single precision can use 16-bit floating point format vectors instead with almost equivalent performance. Summary of the Invention

[0008] According to a first aspect of the present disclosure, a processor is provided, comprising: an extraction circuit, the extraction circuit being used to extract an instruction, the instruction having a field for specifying an opcode and a position of a first source vector including N single-precision elements and a target vector including at least N 16-bit floating-point elements, the opcode instructing an execution circuit to convert each element of the first source vector to a 16-bit floating-point format and store each converted element in a corresponding position of the target vector, the conversion including truncation and rounding when necessary; a decoding circuit, the decoding circuit being used to decode the instruction; and an execution circuit, the execution circuit being used to respond to the instruction according to the opcode.

[0009] According to a second aspect of the present disclosure, a system is provided, comprising a processor and a memory, wherein the processor comprises: an extraction circuit, the extraction circuit being used to extract an instruction, the instruction having a field for specifying an opcode and a position of a first source vector comprising N single-precision elements and a target vector comprising at least N 16-bit floating-point elements, the opcode instructing an execution circuit to convert each element of the first source vector into a 16-bit floating-point format and store each converted element into a corresponding position of the target vector, the conversion including truncation and rounding when necessary; a decoding circuit, the decoding circuit being used to decode the instruction; and an execution circuit, the execution circuit being used to respond to the instruction according to the opcode.

[0010] According to a third aspect of the present disclosure, a method executed by a processor is provided, the method comprising: extracting an instruction using an extraction circuit, the instruction having a field for specifying an opcode and a position of a first source vector including N single-precision elements and a target vector including at least N 16-bit floating-point elements, the opcode instructing an execution circuit to convert each element of the first source vector to a 16-bit floating-point format and store each converted element in a corresponding position of the target vector, the conversion including truncation and rounding when necessary; decoding the instruction using a decoding circuit; and responding to the instruction according to the opcode using the execution circuit.

[0011] According to a fourth aspect of the present disclosure, a machine-readable medium is provided, the machine-readable medium comprising code, which, when executed, causes a machine to perform the method according to the third aspect of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a block diagram illustrating processing components for performing format conversion (or) instructions according to one embodiment;

[0013] Figure 2A is a block diagram illustrating execution of a formatConvert() instruction according to one embodiment;

[0014] Figure 2B is a block diagram illustrating execution of a formatConvert() instruction according to one embodiment;

[0015] Figure 2C is a block diagram illustrating execution of a formatConvert() instruction according to one embodiment;

[0016] Figure 2D is a block diagram illustrating the execution of a 2-input format conversion() instruction according to one embodiment;

[0017] Figure 3Ais pseudo code illustrating an exemplary execution of a formatConvert() instruction according to one embodiment;

[0018] Figure 3B is a pseudo code illustrating an exemplary execution of a 2-input format conversion() instruction according to one embodiment;

[0019] Figure 3C According to one embodiment, a diagram is shown for Figure 3A and 3B Pseudocode of the helper function;

[0020] Figure 4A is a flowchart illustrating a process of a processor responding to a format conversion() instruction according to one embodiment;

[0021] Figure 4B is a flowchart illustrating a process of a processor responding to a 2-input format conversion () instruction according to one embodiment;

[0022] Figure 5A is a block diagram illustrating the format of a Format Convert (VCVTNEPS2BF16) instruction according to one embodiment;

[0023] Figure 5B is a block diagram illustrating the format of a 2-Input Format Convert (VCVTNE2PS2BF16) instruction according to one embodiment;

[0024] Figures 6A-6B is a block diagram illustrating a generic vector friendly instruction format and instruction templates thereof according to some embodiments of the present invention;

[0025] Figure 6A is a block diagram illustrating a generic vector friendly instruction format and its class A instruction templates according to some embodiments of the present invention;

[0026] Figure 6B is a block diagram illustrating a generic vector friendly instruction format and its class B instruction templates according to some embodiments of the present invention;

[0027] Figure 7A is a block diagram illustrating an exemplary specific vector friendly instruction format according to some embodiments of the present invention;

[0028] Figure 7B is a block diagram illustrating the fields of a particular vector friendly instruction format that make up a complete opcode field, according to one embodiment;

[0029] Figure 7C is a block diagram illustrating fields of a particular vector friendly instruction format that make up a register index field, according to one embodiment;

[0030] Figure 7Dis a block diagram illustrating fields of a particular vector friendly instruction format that constitute an enhanced operation field, according to one embodiment;

[0031] Figure 8 is a block diagram of a register architecture according to one embodiment;

[0032] Figure 9A is a block diagram illustrating both an exemplary in-order pipeline and an exemplary register renaming, out-of-order issue / execution pipeline, according to some embodiments;

[0033] Figure 9B is a block diagram illustrating both an exemplary embodiment of an in-order architecture core and an exemplary register renaming, out-of-order issue / execution architecture core to be included in a processor in accordance with some embodiments;

[0034] Figures 10A-10B A block diagram illustrating a more specific exemplary in-order core architecture, which would be one of several logic blocks in a chip (including other cores of the same and / or different types);

[0035] Figure 10A is a block diagram of a single processor core and its connection to an on-chip interconnect network and to its local subset of a level 2 (L2) cache in accordance with some embodiments;

[0036] Figure 10B According to some embodiments Figure 10A An expanded view of a portion of a processor core in ;

[0037] Figure 11 is a block diagram of a processor that may have more than one core, may have an integrated memory controller, and may have integrated graphics according to some embodiments;

[0038] Figure 12-15 is a block diagram of an exemplary computer architecture;

[0039] Figure 12 A block diagram of a system is shown according to some embodiments;

[0040] Figure 13 is a block diagram of a first more specific exemplary system according to some embodiments;

[0041] Figure 14 is a block diagram of a second more specific exemplary system according to some embodiments;

[0042] Figure 15 is a block diagram of a system on a chip (SoC) according to some embodiments; and

[0043] Figure 16is a block diagram contrasted with using a software instruction converter to convert binary instructions in a source instruction set to binary instructions in a target instruction set, according to some embodiments. DETAILED DESCRIPTION

[0044] In the following description, numerous specific details are set forth. However, it is to be understood that some embodiments may be implemented without these specific details. In other cases, well-known circuits, structures, and techniques are not shown in detail to avoid obscuring the understanding of this description.

[0045] References in the specification to "one embodiment," "an embodiment," "an example embodiment," etc. indicate that the described embodiment may include a feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Furthermore, when a feature, structure, or characteristic is described with respect to one embodiment, it is considered to be within the knowledge of those skilled in the art to implement such feature, structure, or characteristic with respect to other embodiments (if explicitly described).

[0046] As mentioned above, applications that process vectors with single-precision sources can also achieve nearly equivalent performance by using 16-bit floating point format vectors instead. Disclosed herein and illustrated by the accompanying figures are Vector Pack Data Format Conversion instructions (VCVTNEPS2BF16 and VCVTNE2PS2BF16) that implement format conversion of one or two source vectors. The VCVTNEPS2BF16 mnemonic indicates: "VCVT" = Vector ConVerT, "NE" = Round to Nearest Even N earest E ven), "PS" = packed single precision source, "2" = to, and "BF16" = BF loat16. The 2-input version of this instruction takes two source vectors, each with N single-precision elements, and produces a destination vector with 2 times N elements in 16-bit floating point format. The 2-input version allows for a balanced solution, where an N-element source vector is converted into an N-element destination vector. Using this balanced solution, all operands, whether source or destination, can be stored in the same type of vector register, whether it is a 128-bit, 256-bit, or 512-bit vector register. At least with reference to Figure 8 An exemplary processor register file is illustrated and described.

[0047] The exposed format conversion (VCVTNEPS2BF16 or VCVTNE2PS2BF16) instructions are expected to achieve comparable quality compared to algorithms using single precision for both source and destination elements, but with reduced memory usage and memory bandwidth requirements, which will help improve performance and power efficiency, especially in machine learning scenarios.

[0048] Related floating-point formats

[0049] The 16-bit floating-point formats used by the disclosed embodiments include bfloat16 (defined by Google Inc. of Mountain View, California), which is sometimes referred to herein as "bf16" or "BF16," and binary16 (promulgated by the Institute of Electrical and Electronics Engineers as IEEE 754-2008), which is sometimes referred to herein as "half precision" or "fp16." The 32-bit floating-point formats used by the disclosed embodiments include binary32 (also promulgated as part of IEEE 754-2008), which is sometimes referred to herein as "single precision" or "fp32."

[0050] Table 1 lists some relevant features and distinctions between related data formats. As shown in the table, all three formats include a sign bit. Binary32, Binary16, and BFloat16 have exponent widths of 8, 5, and 8 bits, respectively, and significand (sometimes referred to herein as "mantissa" or "fraction") bits of 24, 11, and 8, respectively. One advantage of BFloat16 over FP16 is that we can truncate FP32 numbers and still have valid BFloat16 numbers.

[0051] Table 1

[0052] Format Bit symbol index Significant number Binary32 32 1 8-bit 24-bit Binary16 16 1 5-digit 11 people Bfloat16 16 1 8-bit 8-bit

[0053] A processor implementing the disclosed format conversion (VCVTNEPS2BF16 or VCVTNE2PS2BF16) instruction will include extraction circuitry to extract the instruction having fields to specify the opcode and the locations of the first source, second source (for the 2-input version), and destination vectors. Figures 5A-5B, 6A-6B and 7A-7D further illustrate and describe the format of the format conversion (VCVTNEPS2BF16 or VCVTNE2PS2BF16) instruction. The specified source and destination vectors may be located in vector registers or in memory. The opcode instructs the execution circuit to convert each element of the specified source vector to 16-bit floating point, including truncation and rounding (if necessary), and store each converted element in a corresponding position of the specified destination vector. Such a processor will also include: decoding circuitry for decoding the fetched instruction; and execution circuitry for responding to the instruction according to the opcode. At least with reference to Figure 1-2D 、 Figures 9A-9B and Figures 10A-10B The execution circuitry is further described and illustrated below.

[0054] Figure 1 1 is a block diagram illustrating processing components for executing format conversion (or) instructions according to some embodiments. As shown, computing system 100 includes a storage device 101 for storing (one or more) format conversion instructions 103 to be executed. In some embodiments, computing system 100 is a SIMD processor for simultaneously processing multiple elements of a packed data vector.

[0055] In operation, format conversion instruction(s) 103 are fetched from storage device 101 by fetch circuit 105. Format conversion instruction(s) 103 have fields not shown here to specify an opcode and a first source vector comprising N single-precision elements and a location of a destination vector comprising at least N 16-bit floating-point elements, the opcode instructing the execution circuit to convert each element of the specified source vector to 16-bit floating-point format, including truncation and rounding (if necessary), and to store each converted element in a corresponding location of the specified destination vector. Figures 5A-5B , 6A-6B and 7A-7D further illustrate and describe the format conversion (VCVTNEPS2BF16 or VCVTNE2PS2BF16) instruction format.

[0056] The extracted format conversion instruction 107 is decoded by decode circuitry 109, which decodes the extracted format conversion (VCVTNEPS2BF16 or VCVTNE2PS2BF16) instruction 107 into one or more operations. In some embodiments, this decoding includes generating a plurality of micro-operations to be executed by execution circuitry (e.g., execution circuitry 117). Decode circuitry 109 also decodes the instruction suffix and prefix (if used).

[0057] Execution circuitry 117, which has access to register files and memory 115, responds to instruction 111 according to the opcode and is hereinafter referred to as at least Figures 2A-2D, 3A-3C, 4A-4B, 9A-9B and 10A-10B to further describe and illustrate it.

[0058] In some embodiments, register renaming, register allocation, and / or scheduling circuitry 113 provides functionality for one or more of the following: 1) renaming logical operand values ​​to physical operand values ​​(e.g., a register alias table in some embodiments), 2) assigning status bits and flags to decoded instructions, and 3) scheduling decoded format conversion (VCVTNEPS2BF16 or VCVTNE2PS2BF16) instructions 111 for execution on execution circuitry 117 out of the instruction pool (e.g., using a reservation station in some embodiments).

[0059] In some embodiments, write-back circuitry 119 writes back the results of executed instructions. Write-back circuitry 119 and register renaming / scheduling circuitry 113 are optional, as indicated by their dashed borders, in that they may occur at different times, or not at all.

[0060] Figure 2A 1 is a block diagram illustrating the execution of a format conversion () instruction according to one embodiment. As shown in the figure, a computing device 200 (e.g., a processor) receives, extracts, and decodes (the extraction and decoding circuits are not shown here, but at least reference is made to Figure 1 and Figures 9A-9B 1 and 2) format conversion instruction 201. The format conversion instruction 201 includes fields to specify an opcode 202 (VCVTNEPS2BF16) and the location of a first source vector 206 including N single-precision elements and a destination vector 204 including at least N 16-bit floating point (e.g., bfloat16 or binary16) elements.

[0061] Here, N is equal to 4, and the first source 212 and destination 218 vectors specified both have four elements. However, the source and destination vectors are not balanced because they have different widths. Software can issue an unbalanced format conversion instruction 201 by assigning vectors of different sizes to the source and destination vectors, for example, by assigning a 256-bit ymm vector as the source and a 128-bit xmm vector as the destination. Figures 2B-2D The diagram shows a scenario where balancing is achieved by assigning vectors of the same type to both the source and the target. Figure 8 An exemplary processor register file is further illustrated and described below.

[0062] In some embodiments, the format conversion instruction 201 further includes a mask {k} 208 and a zeroing control {z} 210. Figure 5A 、 6A6B and 7A-7D further illustrate and describe the format of the format conversion instruction 201 having an opcode of VCVTNEPS2BF16. Also shown are a designated first source vector 212, execution circuitry 214 including conversion circuitry 216A-D, and a designated destination vector 218.

[0063] In operation, a computing device 200 (e.g., a processor) uses fetch and decode circuitry (not shown) to fetch and decode an instruction 201 having fields to specify an opcode 202 and the locations of first source 206 and destination 204 vectors, the opcode instructing the computing device (e.g., the processor) to convert each element of the specified first source vector 212 to a 16-bit floating point format (e.g., bfloat16), including truncation and rounding (if necessary) by converter circuits 216A-D, and to store each converted element in a corresponding location of the specified destination vector 218. As at least referenced herein, Figure 5A 、 6A As further illustrated and described in FIG6B and FIG7A-7D , the instruction 201 may specify a different vector length in other embodiments, such as 128 bits, 512 bits, or 1024 bits. Execution circuitry 214 responds to the instruction based on the opcode 202 .

[0064] Figure 2B 1 is a block diagram illustrating the execution of a format conversion () instruction according to one embodiment. As shown in the figure, the computing device 220 (e.g., a processor) receives, extracts and decodes (the extraction and decoding circuits are not shown here, but at least reference is made to Figure 1 and Figures 9A-9B 16) format conversion instruction 221 that includes fields to specify an opcode 222 (VCVTNEPS2BF16) and the location of a first source vector 226 including N single-precision elements and a destination vector 224 including at least N 16-bit floating point (e.g., bfloat16 or binary16) elements.

[0065] Here, balancing is achieved by assigning registers of the same type as the designated first source 232 and destination 238. However, the designated destination vector 238, which is half the width of the designated first source vector 232, has twice as many entries. In operation, the converted entries are written to the first four destination entries, and zeros are written to the remaining four entries.

[0066] In some embodiments, the format conversion instruction 221 also includes a mask {k} 228 and a zeroing control {z} 230. Figure 5A 、 6A6B and 7A-7D further illustrate and describe the format of the format conversion instruction 221 having an opcode of VCVTNEPS2BF16. Also shown are a designated first source vector 232, execution circuitry 234, and a designated destination vector 238, wherein execution circuitry 214 includes conversion circuitry 236A-D.

[0067] In operation, the computing device 220 (e.g., a processor) uses fetch and decode circuitry (not shown) to fetch and decode instruction 221, which has fields to specify an opcode 222 (i.e., VCVTNEPS2BF16) and the locations of the first source 226 and destination 224 vectors. The opcode instructs the computing device 220 (e.g., a processor) to use converters 236A-D in the execution circuitry 234 to convert each element of the specified first source vector 232 to a 16-bit floating point format (e.g., bfloat16), including truncation and rounding (if necessary), and to store each converted element in a corresponding position of the specified destination vector 218. Here, the corresponding destination vector positions include the first four elements, with zeros written to the remaining four elements. As at least with reference to Figure 5A 、 6A As further illustrated and described in FIG6B and FIG7A-7D , the instruction 221 may specify a different vector length in other embodiments, such as 128 bits, 512 bits, or 1024 bits. Execution circuitry 234 responds to the instruction based on the opcode 202 .

[0068] Figure 2C 1 is a block diagram illustrating the execution of a format conversion () instruction according to one embodiment. As shown in the figure, the computing device 240 (e.g., a processor) receives, extracts and decodes (the extraction and decoding circuits are not shown here, but at least reference is made to Figure 1 and Figures 9A-9B 16) format conversion instruction 241 that includes fields to specify an opcode 242 (VCVTNEPS2BF16) and the location of a first source vector 246 including N single-precision elements and a destination vector 244 including at least N 16-bit floating point (e.g., bfloat16 or binary16) elements.

[0069] Here, balancing is achieved by assigning registers of the same type as the first source 252 and destination 258 vectors specified. However, the destination vector 258, which is half the width of the first source vector 252 specified, has twice as many entries. In operation, the converted entries are written to the first four destination entries, and zeros are written to the remaining four entries. Figure 2C This is not shown, but will be done implicitly in this embodiment.

[0070] Implicit zeroing is the default treatment for masked elements in some embodiments. In other embodiments, an architectural model-specific register (MSR) is programmed by software to control whether zeroing or masking is applied to masked elements. In still other embodiments, the zeroing behavior is specified by a format conversion instruction.

[0071] In some embodiments, the format conversion instruction 241 also includes a mask {k} 248 and a zeroing control {z} 250. Figure 5A 、 6A -6B and 7A-7D further illustrate and describe the format of the format conversion instruction 241 having an opcode of VCVTNEPS2BF16.

[0072] Also shown are a designated first source vector 252, execution circuitry 254, and a designated destination vector 258, where execution circuitry 214 includes conversion circuitry 256A-D.

[0073] In operation, a computing device 240 (e.g., a processor) uses fetch and decode circuitry (not shown) to fetch and decode an instruction 241 having fields specifying an opcode 242 (i.e., VCVTNEPS2BF16) and the locations of first source 246 and destination 244 vectors. The opcode instructs the computing device 240 (e.g., the processor) to use converters 256A-D in execution circuitry 254 to convert each element of the specified first source vector 252 to a 16-bit floating point format (e.g., bfloat16). Converter circuitry 256A-D includes truncation and rounding (if necessary) and stores each converted element in a corresponding position of the specified destination vector 258. Here, the corresponding destination vector positions include the first four elements, with zeros implicitly written to the remaining four elements. Implicit zeroing is the default handling of masked elements in some embodiments. In other embodiments, a model-specific register (MSR) on the architecture is programmed by software to control whether zeroing or masking is applied to masked elements. In yet other embodiments, the zeroing behavior is specified by a format conversion instruction.

[0074] As at least reference Figure 5A 、 6A As further illustrated and described in FIG6B and FIG7A-7D , the instruction 241 may specify a different vector length in other embodiments, such as 128 bits, 512 bits, or 1024 bits. The execution circuit 254 responds to the instruction according to the operation code 242 .

[0075] Figure 2D1 is a block diagram illustrating the execution of a format conversion (VCVTNE2PS2BF16) instruction according to one embodiment. As shown, a computing device 260 (e.g., a processor) receives, extracts, and decodes (the extraction and decoding circuitry is not shown here, but at least reference is made to Figure 1 and Figures 9A-9B 2 (not shown and described herein) format conversion instruction 261 includes fields to specify an opcode 262 (VCVTNE2PS2BF16) and the location of first and second source vectors 266 and 268 comprising N single-precision elements and a destination vector 264 comprising at least N 16-bit floating point (e.g., bfloat16 or binary16) elements. Here, N is equal to 4 and the specified destination vector 264 includes 8 elements.

[0076] Here, the width of the destination vector is half that of the source vector, but balance is achieved by assigning two source vectors whose elements are to be converted and written to the destination. In operation, the converted entries from the designated first source 272A are written to the first four entries of the designated destination 278, and the converted entries from the designated second source 272B are written to the last four entries of the designated destination 278.

[0077] In some embodiments, the format conversion instruction 261 also includes a mask {k} 268 and a zeroing control {z} 270. Figure 5A 、 6A -6B and 7A-7D further illustrate and describe the format of the format conversion instruction 261 having an opcode of VCVTNEPS2BF16.

[0078] Also shown are designated first and second source vectors 272A-B, execution circuitry 274 including conversion circuitry 276A-H, and designated destination vector 278.

[0079] In operation, a computing device 260 (e.g., a processor) uses fetch and decode circuitry (not shown) to fetch and decode instruction 261, which has fields that specify an opcode 262 (i.e., VCVTNE2PS2BF16) and the locations of first and second source 266 and 268 and destination 264 vectors. The opcode instructs the computing device 260 (e.g., the processor) to use converters 276A-H in execution circuitry 274 to convert each element of the specified first and second source vectors 272A-B to a 16-bit floating point format (e.g., bfloat16), including truncation and rounding (if necessary), and to store each converted element in a corresponding location of the specified destination vector 278. Here, the first four elements of the specified destination 278 correspond to the specified first source 272A, and the last four elements of the specified destination 278 correspond to the specified second source 272B. As at least with reference to FIG. Figure 5A 、 6A As further illustrated and described in FIG6B and FIG7A-7D , the instruction 261 may specify a different vector length in other embodiments, such as 128 bits, 512 bits, or 1024 bits. Execution circuitry 274 responds to the instruction according to the opcode 262 .

[0080] Figure 3A is pseudo code illustrating an exemplary execution of a format conversion (VCVTNEPS2BF16) instruction according to one embodiment. As shown, the format conversion instruction 301 has fields to specify an opcode 302 (VCVTNEPS2BF16) and the locations of the first source 306 (src) and destination 304 (dest) vectors, which can be any of 128 bits, 256 bits, and 512 bits, according to a constant VL instantiated in the code and representing "vector length". In some embodiments, the instruction 301 also has fields to specify a mask 308 and a zeroing control 310. The pseudo code 315 also illustrates the use of a write mask to control whether each destination element is masked, masked elements are zeroed, or merged (as at least with reference to the vector length). Figure 5A 、 6A -6B and 7A-7D as further illustrated and described, some embodiments of the format conversion instruction include fields to specify the mask and control whether to zero or merge). At least with reference to Figures 2A-2C , 4A and 9A-9B further illustrate and describe the execution of the format conversion instruction 301.

[0081] Figure 3Bis pseudo code illustrating an exemplary execution of a 2-input format conversion () instruction according to one embodiment. As shown, the format conversion instruction 321 has fields to specify an opcode 322 (VCVTNE2PS2BF16) and the locations of the first source 326 (src1), the second source 328 (src2), and the destination 324 (dest) vectors. The destination vector can be any one of 128 bits, 256 bits, and 512 bits according to the constant VL. Here, the source vector location can be in memory or in a register. In some embodiments, the format conversion instruction 321 has fields to specify a write mask {k} 330 and a zeroing control {z} 331. Pseudo code 335 also shows the use of a write mask to control whether each destination element is masked, the masked elements are zeroed, or merged (as at least with reference to Figure 5A 、 6A -6B and 7A-7D as further illustrated and described, some embodiments of the format conversion instruction include fields to specify the mask and control whether to zero or merge). At least with reference to Figure 2D 、 4B 9A-9B further illustrate and describe the execution of the format conversion instruction 321.

[0082] Figure 3C According to one embodiment, a diagram is shown for Figures 3A-3B Here, pseudocode 354 defines a helper function, convert_fp32_to_bfloat16(), which converts from binary32 format to bfloat16 format.

[0083] Pseudocode 340 illustrates that the disclosed embodiment, unlike a simple conversion that would simply truncate the lower sixteen bits of a binary32 number, advantageously performs rounding of normal numbers and accounts for rounding bias. The code illustrates that the format conversion instruction has improved rounding behavior compared to simply truncating. The rounding behavior of the disclosed embodiment promotes more accurate calculations than conversions performed by truncation. In some embodiments, the execution circuitry adheres to rounding behavior according to the rounding rules promulgated as IEEE 754, such as "NE" indicating rounding to the nearest even. In some embodiments, the rounding behavior is specified by the instruction, such as by including the suffix "NE" in the opcode to indicate rounding to the nearest even. In other embodiments, the rounding behavior uses a default behavior such as "NE". In still other embodiments, the rounding behavior is controlled by software-configured architectural model-specific registers (MSRs).

[0084] Pseudo-code 340 also illustrates that the disclosed embodiments perform truncation when necessary, such as when an input to a function is not a number (nan).

[0085] At least for reference Figures 2A-2D, 3A-3B, 4A-4B and 9A-9B further illustrate and describe the execution of the format conversion instruction.

[0086] Figure 4A 1 is a process flow diagram illustrating a processor responding to a format-convert() instruction according to one embodiment. The format-convert instruction 401 includes fields to specify an opcode 402 (VCVTNEPS2BF16) and the location of a first source vector 406 including N single-precision elements and a destination vector 404 including at least N 16-bit floating point (e.g., bfloat16 or binary16) elements.

[0087] As shown, the processor responds to a decoded format conversion instruction by executing flow 400. At 421, the processor uses extraction circuitry to extract an instruction having fields specifying an opcode (e.g., VCVTNEPS2BF16) and a first source vector comprising N single-precision elements and a destination vector comprising at least N 16-bit floating-point (e.g., bfloat16 or binary16) elements. The opcode instructs the execution circuitry to convert each element of the specified source vector to 16-bit floating point, including truncation and rounding (if necessary), and to store each converted element in a corresponding location of the specified destination vector. At 423, the processor uses decoding circuitry to decode the extracted instruction. In some embodiments, the processor schedules execution of the decoded instruction at 425. At 427, the processor uses execution circuitry to respond to the instruction based on the opcode. In some embodiments, the processor commits the results of the executed instruction at 429. Operations 425 and 429 are optional, as indicated by their dashed borders, because they can occur at different times or not at all.

[0088] Figure 4B FIG2 is a flow diagram illustrating a process of a processor responding to a 2-input format conversion () instruction according to one embodiment. The format conversion instruction 451 includes fields to specify an opcode 452 (VCVTNE2PS2BF16) and the locations of first and second source vectors 456 and 462 comprising N single-precision elements and a destination vector 454 comprising at least N 16-bit floating point (e.g., bfloat16 or binary16) elements.

[0089] Note that the present invention is not intended to be limited to any particular mnemonic for the opcode. Here, VCVTNEPS2BF16 is chosen as a mnemonic with letters representing various instruction characteristics. For example, "VCVT" is chosen to indicate vector conversion ( V ector C on V er TFor another example, "NE" is chosen to indicate that the rounding mode prescribed by IEEE 754 (here, nearest even) is selected. "2PS" means 2-pack single. "2" means "to". Finally, "BF16" means bfloat16.

[0090] As shown, the processor responds to a decoded format conversion instruction by executing flow 450. At 471, the processor uses fetch circuitry to fetch an instruction having fields specifying an opcode (e.g., VCVTNE2PS2BF16) and the location of a first and second source vector comprising N single-precision elements and a destination vector comprising at least N 16-bit floating-point (e.g., bfloat16 or binary16) elements. The opcode instructs the execution circuitry to convert each element of the specified first and second source vectors to 16-bit floating point, including truncation and rounding (if necessary), and to store each converted element in a corresponding location of the specified destination vector. At 473, the processor uses decode circuitry to decode the fetched instruction. In some embodiments, the processor schedules execution of the decoded instruction at 475. At 477, the processor uses the execution circuitry to respond to the instruction based on the opcode. In some embodiments, the processor commits the results of the executed instruction at 479. Operations 475 and 479 are optional, as indicated by their dashed borders, because they can occur at different times or not at all.

[0091] Figure 5A 1 is a block diagram illustrating the format of a Format Convert (VCVTNEPS2BF16) instruction according to one embodiment. As shown, the Format Convert instruction 500 includes fields for specifying an opcode 502 (VCVTNEPS2BF16) and the locations of a target 504 and a first source 506 vector. The source and target vectors may each be located in a register or in memory.

[0092] The opcode 502 is shown as including an asterisk, which indicates that various optional fields can be added to the opcode as a prefix or suffix. That is, the format conversion instruction 500 also includes optional parameters to affect the instruction behavior, including mask {k} 508, zeroing control {z} 510, element format 514, vector size (N) 516, and rounding mode 518. One or more of the instruction modifiers 508, 510, 514, 516, and 518 can be specified using a prefix or suffix of the opcode 502.

[0093] In some embodiments, one or more of the optional instruction modifiers 508, 510, 514, 516, and 518 are encoded in an immediate field (not shown) that is optionally included with the instruction 500. In some embodiments, one or more of the optional instruction modifiers 508, 510, 514, 516, and 518 are specified via a configuration register, such as a model specific register (MSR) included in the instruction set architecture.

[0094] At least for reference Figure 5B 、 6A -6B and 7A-7D further illustrate and describe the format of the format conversion instruction 500.

[0095] Figure 5B FIG2 is a block diagram illustrating the format of a 2-input format conversion (VCVTNE2PS2BF16) instruction according to one embodiment. As shown, the format conversion instruction 550 includes fields for specifying an opcode 552 (VCVTNE2PS2BF16) and the locations of a destination 554, a first source 556, and a second source 552 vector. The source and destination vectors can each be located in a register or in memory.

[0096] The opcode 552 is shown as including an asterisk, which indicates that various optional fields can be added to the opcode as a prefix or suffix. That is, the format conversion instruction 550 also includes optional parameters to affect the instruction behavior, including mask {k} 558, zeroing control {z} 560, element format 564, vector size (N) 566, and rounding mode 568. One or more of the instruction modifiers 558, 560, 564, and 566 can be specified using a prefix or suffix of the opcode 552.

[0097] In some embodiments, one or more of the optional instruction modifiers 558, 560, 564, 566, and 568 are encoded in an immediate field (not shown) that is optionally included with the instruction 550. In some embodiments, one or more of the optional instruction modifiers 558, 560, 564, 566, and 568 are specified via a configuration register, such as a model specific register (MSR) included in the instruction set architecture.

[0098] At least for reference Figure 5A 、 6A -6B and 7A-7D further illustrate and describe the format of the format conversion instruction 550.

[0099] instruction set

[0100] An instruction set may include one or more instruction formats. A given instruction format may define various fields (e.g., the number of bits, the position of the bits) to specify the operation to be performed (e.g., an opcode) and the (one or more) operands and / or other (one or more) data fields (e.g., a mask) on which the operation is to be performed, etc. Some instruction formats are further decomposed by the definition of instruction templates (or subformats). For example, the instruction templates of a given instruction format may be defined as having different subsets of the fields of the instruction format (the included fields are generally in the same order, but at least some may have different bit positions because there are fewer included fields) and / or may be defined as interpreting given fields differently. Thus, each instruction of the ISA is expressed using a given instruction format (and, if defined, expressed in a given one of the instruction templates of the instruction format) and includes fields for specifying the operation and the operand. For example, an exemplary ADD instruction has a specific opcode and instruction format that includes an opcode field to specify the opcode and an operand field to select the operand (source 1 / destination and source 2); and an appearance of this ADD instruction in an instruction stream will have specific contents in the operand field that selects the specific operand. A set of SIMD extensions known as Advanced Vector Extensions (AVX) (AVX1 and AVX2) and using the Vector Extension (VEX) encoding scheme have been released and / or published (e.g., see 64 and IA-32 Architectures Software Developer's Manual, September 2014; and see Advanced Vector Extensions Programming Reference, October 2014).

[0101] Example instruction format

[0102] The embodiments of the instruction(s) described herein may be implemented in different formats. In addition, exemplary systems, architectures, and pipelines are described in detail below. The embodiments of the instruction(s) may be executed on such systems, architectures, and pipelines, but are not limited to those described in detail.

[0103] Generic vector-friendly instruction format

[0104] The vector friendly instruction format is an instruction format suitable for vector instructions (e.g., has certain fields specific to vector operations). Although embodiments are described in which both vector and scalar operations are supported by the vector friendly instruction format, alternative embodiments use only the vector friendly instruction format for vector operations.

[0105] Figures 6A-6B is a block diagram illustrating a generic vector friendly instruction format and instruction templates thereof according to some embodiments of the present invention. Figure 6A is a block diagram illustrating a generic vector friendly instruction format and its class A instruction templates according to some embodiments of the present invention; and Figure 6B 6 is a block diagram illustrating a generic vector friendly instruction format and its class B instruction templates according to some embodiments of the present invention. Specifically, class A and class B instruction templates are defined for the generic vector friendly instruction format 600, both of which include a no memory access 605 instruction template and a memory access 620 instruction template. The term "generic" in the context of the vector friendly instruction format means that the instruction format is not tied to any specific instruction set.

[0106] Although embodiments of the present invention will be described in which the vector friendly instruction format supports: 64-byte vector operand lengths (or sizes) with either 32-bit (4-byte) or 64-bit (8-byte) data element widths (or sizes) (thus, a 64-byte vector consists of 16 doubleword-sized elements or 8 quadword-sized elements); 64-byte vector operand lengths (or sizes) with either 16-bit (2-byte) or 8-bit (1-byte) data element widths (or sizes); 32-byte vector operand lengths (or sizes) with either 32-bit (4-byte), 64-bit (8-byte) data element widths (or sizes); 4-bit (8-byte), 16-bit (2-byte), or 8-bit (1-byte) data element width (or size); and 16-byte vector operand length (or size) with 32-bit (4-byte), 64-bit (8-byte), 16-bit (2-byte), or 8-bit (1-byte) data element width (or size); although alternative embodiments may support more, fewer, and / or different vector operand sizes (e.g., 256-byte vector operands) with more, fewer, or different data element widths (e.g., 128-bit (16-byte) data element width).

[0107] Figure 6A The class A instruction templates in include: 1) within the no memory access 605 instruction template, the no memory access, full rounding control type operation 610 instruction template and the no memory access, data transformation type operation 615 instruction template are shown; and 2) within the memory access 620 instruction template, the memory access, transient 625 instruction template and the memory access, non-transient 630 instruction template are shown. Figure 6B The category B instruction templates include: 1) within the no memory access 605 instruction template, the no memory access, write mask control, partial rounding control type operation 612 instruction template and the no memory access, write mask control, vsize type operation 617 instruction template are shown; and 2) within the memory access 620 instruction template, the memory access, write mask control 627 instruction template is shown.

[0108] The general vector friendly instruction format 600 includes Figures 6A-6BThe following fields are listed below in the order shown.

[0109] Format field 640 - The specific value in this field (the instruction format identifier value) uniquely identifies the vector friendly instruction format, and thus identifies the occurrence of instructions in the vector friendly instruction format in the instruction stream. As such, this field is optional, as it is not required for instruction sets that only have the general vector friendly instruction format.

[0110] Basic operation field 642 - its content distinguishes different basic operations.

[0111] Register index field 644 - its contents specify the location of the source and target operands, either in registers or in memory, either directly or through address generation. These include a sufficient number of bits to select N registers from a PxQ (e.g., 32x512, 16x128, 32x1024, 64x1024) register file. While in one embodiment N may be up to three sources and one target register, alternative embodiments may support more or fewer source and target registers (e.g., may support up to two sources, one of which also serves as a target, may support up to three sources, one of which also serves as a target, may support up to two sources and one target).

[0112] Modifier field 646 - its content distinguishes occurrences of instructions in the generic vector instruction format that specify memory access from those that do not; that is, distinguishes between the no memory access 605 instruction template and the memory access 620 instruction template. Memory access operations read and / or write to the memory hierarchy (in some cases using values ​​in registers to specify the source and / or destination addresses), while non-memory access operations do not read and / or write to the memory hierarchy (e.g., the source and destination are registers). While in one embodiment this field also selects between three different ways to perform memory address calculations, alternative embodiments may support more, fewer, or different ways to perform memory address calculations.

[0113] Enhanced Operation Field 650 - Its contents distinguish which of a variety of different operations is to be performed in addition to the basic operation. This field is context-dependent. In some embodiments, this field is divided into a category field 668, an alpha field 652, and a beta field 654. The enhanced operation field 650 allows a common group of operations to be performed in a single instruction instead of two, three, or four instructions.

[0114] Scale field 660 - its content allows scaling the contents of the index field for memory address generation (e.g., for memory addresses using 2 缩放比例 *Address generation of index+base).

[0115] Displacement field 662A - its contents are used as part of memory address generation (e.g., for memory addresses using 2 缩放比例 *Address generation of index+base+displacement).

[0116] Displacement Factor Field 662B (note that listing Displacement Field 662A immediately above Displacement Factor Field 662B indicates that one or the other is being used) - its contents are used as part of address generation; it specifies the displacement factor to be scaled by the size (N) of the memory access - where N is the number of bytes in the memory access (e.g., for a memory access using 2 缩放比例 * address generation using index + base + scaled displacement). Redundant low-order bits are ignored, and therefore, the contents of the displacement factor field are multiplied by the total size of the memory operand (N) to generate the final displacement to be used in calculating the effective address. The value of N is determined by the processor hardware at runtime based on the full opcode field 674 (described later herein) and the data manipulation field 654C. The displacement field 662A and the displacement factor field 662B are optional in that they are not used for the no memory access 605 instruction template, and / or different embodiments may implement only one or neither of them.

[0117] Data element width field 664 - its contents distinguish which of several data element widths is used (in some embodiments for all instructions; in other embodiments for only some of the instructions). This field is optional, as it is not needed if only one data element width is supported and / or some aspect of the opcode is used to support the data element width.

[0118] Write Mask Field 670 – Its contents control, on a per-data-element basis, whether that data element position in the target vector operand reflects the results of the base and enhanced operations. Class A instruction templates support merge-write masking, while class B instruction templates support both merge-write masking and zero-write masking. When merged, the vector mask allows any set of elements in the target to be protected from updates during the execution of any operation (specified by the base and enhanced operations); in one embodiment, the old value of each element of the target whose corresponding mask bit has a value of 0 is preserved. Separately, the zero-vector mask allows any set of elements in the target to be zeroed during the execution of any operation (specified by the base and enhanced operations); in one embodiment, elements in the target are set to 0 if their corresponding mask bit has a value of 0. A subset of this functionality is the ability to control the vector length of the operation being performed (i.e., the stride of modified elements, from first to last); however, the modified elements do not have to be contiguous. Thus, the write mask field 670 enables some vector operations, including loads, stores, arithmetic, logical, and more. Although embodiments of the present invention have been described in which the contents of write mask field 670 select one of several write mask registers containing the write mask to be used (such that the contents of write mask field 670 indirectly identify the masking to be performed), alternative embodiments alternatively or additionally allow the contents of mask write field 670 to directly specify the masking to be performed.

[0119] Immediate field 672 - its content allows specification of an immediate. This field is optional in that it is not present in implementations that do not support the generic vector friendly format for immediates and it is not present in instructions that do not use immediates.

[0120] Category field 668 - its content distinguishes different categories of instructions. Figures 6A-6B , the content of this field selects between Category A and Category B instructions. Figures 6A-6B In , rounded squares are used to indicate that a specific value exists in a field (e.g., Figures 6A-6B Category A 668A and Category B 668B for category field 668, respectively).

[0121] Instruction template for category A

[0122] In the case of the non-memory access 605 instruction templates of class A, the alpha field 652 is interpreted as the RS field 652A, the contents of which distinguish which of the different enhanced operation types is to be performed (e.g., rounding 652A.1 and data transformation 652A.2 are specified for the no-memory-access-round-type operation 610 and no-memory-access-data-transformation-type operation 615 instruction templates, respectively), while the beta field 654 distinguishes which of the specified types of operations is to be performed. In the no-memory-access 605 instruction templates, the scale field 660, the displacement field 662A, and the displacement scale field 662B are not present.

[0123] No memory access instruction templates - full rounding control type operations

[0124] In the no memory access full round control type operation 610 instruction template, the beta field 654 is interpreted as a round control field 654A, whose content(s) provide static rounding. Although in the described embodiment of the present invention, the round control field 654A includes a suppress all floating-point exceptions (SAE) field 656 and a round operation control field 658, alternative embodiments may support encoding both of these concepts into the same field or may have only one or the other of these concepts / fields (e.g., may have only the round operation control field 658).

[0125] SAE field 656 - its content distinguishes whether exception event reporting is disabled; when the content of the SAE field 656 indicates that suppression is enabled, the given instruction does not report any kind of floating point exception flags and does not raise any floating point exception handler.

[0126] Round operation control field 658 - its contents distinguish which of a set of rounding operations is to be performed (e.g., round up, round down, round towards zero, and round towards nearest). Thus, the round operation control field 658 allows the rounding mode to be changed on a per-instruction basis. In some embodiments where the processor includes a control register for specifying the rounding mode, the contents of the round operation control field 650 override the register value.

[0127] No memory access instruction template - data transformation operation

[0128] In the no memory access data transform type operation 615 instruction template, the beta field 654 is interpreted as a data transform field 654B, whose content distinguishes which of several data transforms is to be performed (eg, no data transform, swizzle, broadcast).

[0129] In the case of a memory access 620 instruction template of class A, the alpha field 652 is interpreted as an eviction hint field 652B, the content of which distinguishes which of the eviction hints is to be used (in Figure 6A 652B.1 and non-transitory 652B.2 are specified for the memory access transient 625 instruction template and the memory access non-transitory 630 instruction template, respectively, while the beta field 654 is interpreted as a data manipulation field 654C, the contents of which distinguish which of several data manipulation operations (also called primitives) is to be performed (e.g., no manipulation; broadcast; upcast of source; and downcast of destination). The memory access 620 instruction template includes a scale field 660 and optionally a displacement field 662A or a displacement factor field 662B.

[0130] Vector memory instructions perform vector loads from memory and vector stores to memory, with support for conversions. Like regular vector instructions, vector memory instructions transfer data to / from memory in a per-data-element manner, where the actual elements transferred are specified by the contents of the vector mask selected as the write mask.

[0131] Memory Access Instruction Templates - Transient

[0132] Transient data is data that is likely to be reused soon enough to benefit from caching. However, this is a hint, and different processors can implement it in different ways, including ignoring the hint entirely.

[0133] Memory access instruction templates - non-transient

[0134] Non-transient data is data that is unlikely to be reused quickly enough to benefit from caching in level 1 cache and should be given eviction priority. However, this is a hint, and different processors may implement it in different ways, including ignoring the hint entirely.

[0135] Instruction template for category B

[0136] In the case of class B instruction templates, the alpha field 652 is interpreted as a write mask control (Z) field 652C, the content of which distinguishes whether the write masking controlled by the write mask field 670 should be merge or zero.

[0137] In the case of the non-memory access 605 instruction templates of class B, a portion of the beta field 654 is interpreted as the RL field 657A, the contents of which distinguish which of the different enhanced operation types is to be performed (e.g., rounding 657A.1 and vector length (VSIZE) 57A.2 are specified for the no memory access, write mask control, partial rounding control type operation 612 instruction template and the no memory access, write mask control, VSIZE type operation 617 instruction template, respectively), while the remainder of the beta field 654 distinguishes which of the specified types of operations is to be performed. In the no memory access 605 instruction templates, the scale field 660, the displacement field 662A, and the displacement scale field 662B are not present.

[0138] In the no memory access, write mask control, partial round control type operation 610 instruction template, the remainder of the beta field 654 is interpreted as the round operation field 659A and exception event reporting is disabled (the given instruction does not report any kind of floating point exception flags and does not trigger any floating point exception handler).

[0139] Round Operation Control Field 659A - As with the round operation control field 658, its contents distinguish which of a set of rounding operations is to be performed (e.g., round up, round down, round toward zero, and round toward nearest). Thus, the round operation control field 659A allows the rounding mode to be changed on a per-instruction basis. In some embodiments where the processor includes a control register for specifying the rounding mode, the contents of the round operation control field 659A override the register value.

[0140] In the no memory access, write mask control, VSIZE type operation 617 instruction template, the remainder of the beta field 654 is interpreted as a vector length field 659B, the contents of which distinguish which of several data vector lengths to operate on (e.g., 128, 256, or 512 bytes).

[0141] In the case of a memory access 620 instruction template of class B, a portion of the beta field 654 is interpreted as a broadcast field 657B, the contents of which distinguish whether a broadcast-type data operation is to be performed, while the remainder of the beta field 654 is interpreted as a vector length field 659B. The memory access 620 instruction template includes a scaling field 660 and optionally a displacement field 662A or a displacement factor field 662B.

[0142] For the generic vector friendly instruction format 600, the full opcode field 674 is shown as including the format field 640, the basic operation field 642, and the data element width field 664. Although one embodiment is shown in which the full opcode field 674 includes all of these fields, the full opcode field 674 may include only some of these fields in embodiments that do not support all of these fields. The full opcode field 674 provides an operation code (opcode).

[0143] The enhanced operation field 650, the data element width field 664, and the write mask field 670 allow these features to be specified on a per-instruction basis in the generic vector friendly instruction format.

[0144] The combination of writing the mask field and the data element width field creates typed instructions because they allow the mask to be applied based on different data element widths.

[0145] The various instruction templates found within categories A and B are beneficial in different situations. In some embodiments of the present invention, different processors or different cores within a processor may support only category A, only category B, or both categories. For example, a high-performance general-purpose out-of-order core intended for general-purpose computing may support only category B, a core intended primarily for graphics and / or scientific (throughput) computing may support only category A, and a core intended for both may support both (of course, cores with some mix of templates and instructions from both categories, but not all templates and instructions from both categories, are within the scope of the present invention). In addition, a single processor may include multiple cores, all of which support the same category or where different cores support different categories. For example, in a processor with separate graphics and general-purpose cores, one of the graphics cores intended primarily for graphics and / or scientific computing may support only category A, while one or more of the general-purpose cores may be a high-performance general-purpose core intended for general-purpose computing with out-of-order execution and register renaming that supports only category B. Another processor without a separate graphics core may include one or more general-purpose in-order or out-of-order cores that support both categories A and B. Of course, in different embodiments of the present invention, features from one category may also be implemented in another category. A program written in a high-level language will be placed (e.g., just-in-time or statically compiled into) a variety of different executable forms, including: 1) a form having only instructions from the class(es) supported by the target processor for execution; or 2) a form having alternative routines written using different combinations of instructions from all classes and having control flow code that selects the routine to be executed based on the instructions supported by the processor currently executing the code.

[0146] Exemplary specific vector friendly instruction format

[0147] Figure 7A is a block diagram illustrating an exemplary specific vector friendly instruction format according to some embodiments of the invention. Figure 7A A specific vector friendly instruction format 700 is shown that is specific in the sense that it specifies the location, size, interpretation, and order of fields, as well as the values ​​of some of these fields. The specific vector friendly instruction format 700 can be used to extend the x86 instruction set so that some of the fields are similar or identical to those used in the existing x86 instruction set and its extensions (e.g., AVX). This format is consistent with the prefix encoding field, real opcode byte field, MOD R / M field, SIB field, displacement field, and immediate field of the existing x86 instruction set with extensions. The figure shows the vector friendly instruction format 700 from Figure 7A The fields from Figure 6 are mapped to the fields from Figure 6.

[0148] It should be understood that while embodiments of the present invention are described with reference to the specific vector friendly instruction format 700 in the context of the generic vector friendly instruction format 600 for illustrative purposes, the present invention is not limited to the specific vector friendly instruction format 700 unless otherwise stated. For example, the generic vector friendly instruction format 600 contemplates a variety of possible sizes for various fields, while the specific vector friendly instruction format 700 is illustrated as having fields of specific sizes. As a specific example, while the data element width field 664 is illustrated as a one-bit field in the specific vector friendly instruction format 700, the present invention is not limited thereto (that is, the generic vector friendly instruction format 600 contemplates other sizes for the data element width field 664).

[0149] The general vector friendly instruction format 600 includes Figure 7A The following fields are listed below in the order shown.

[0150] EVEX prefix (bytes 0-3) 702 - encoded in four-byte form.

[0151] Format field 640 (EVEX byte 0, bits [7:0]) - The first byte (EVEX byte 0) is the format field 640 and it contains 0x62 (a unique value used in some embodiments to distinguish the vector friendly instruction format).

[0152] The second through fourth bytes (EVEX bytes 1-3) include several bit fields that provide specific capabilities.

[0153] REX field 705 (EVEX byte 1, bits [7-5]) - consists of the EVEX.R bit field (EVEX byte 1, bits [7]-R), the EVEX.X bit field (EVEX byte 1, bits [6]-X), and the 657BEX byte 1, bits [5]-B). The EVEX.R, EVEX.X, and EVEX.B bit fields provide the same functionality as the corresponding VEX bit fields and are encoded using inverted form, i.e., ZMM0 is encoded as 1111B and ZMM15 is encoded as 0000B. The other fields of the instruction encode the lower three bits of the register index as is known in the art, such that Rrrr, Xxxx, and Bbbb can be formed by adding EVEX.R, EVEX.X, and EVEX.B.

[0154] REX' 710A - This is the first portion of the REX' field 710 and is the EVEX.R' bit field (EVEX byte 1, bit [4] - R') used to encode the upper 16 or lower 16 of the extended 32 register set. In some embodiments, this bit, along with the other bits shown below, are stored in bit-reversed format to distinguish it from the BOUND instruction (in the well-known x86 32-bit mode), which has a true opcode byte of 62 but does not accept a value of 11 in the MOD field in the MOD R / M field (described below); alternative embodiments of the present invention do not store this bit and the other bits indicated below in reversed format. A value of 1 is used to encode the lower 16 registers. In other words, R'Rrrr is formed by combining EVEX.R', EVEX.R, and the other RRRs from other fields.

[0155] Opcode map field 715 (EVEX byte 1, bits [3:0] - mmmm) - its contents encode the implied dominant opcode type (0F, 0F 38, or 0F 3).

[0156] Data element width field 664 (EVEX byte 2, bit [7] - W) - represented by the symbol EVEX.W. EVEX.W is used to define the granularity (size) of the data type (32-bit data elements or 64-bit data elements).

[0157] EVEX.vvvv 720 (EVEX byte 2, bits [6:3] - vvvv) - The effects of EVEX.vvvv can include the following: 1) EVEX.vvvv encodes the first source register operand specified in inverted (one's complement) form and is valid for instructions with two or more source operands; 2) EVEX.vvvv encodes the destination register operand specified in inverted (one's complement) form for certain vector shifts; or 3) EVEX.vvvv does not encode any operands; this field is reserved and should contain 1111b. Thus, EVEX.vvvv field 720 encodes the four low-order bits of the first source register designator stored in inverted (one's complement) form. Depending on the instruction, an additional different EVEX bit field is used to extend the designator size to 32 registers.

[0158] EVEX.U 668 Class field (EVEX byte 2, bit [2] - U) - If EVEX.U = 0, it indicates Class A or EVEX.U0; if EVEX.U = 1, it indicates Class B or EVEX.U1.

[0159] Prefix encoding field 725 (EVEX byte 2, bits [1:0]-pp) – provides additional bits for the base operation field. In addition to providing support for legacy SSE instructions in the EVEX prefix format, this also has the benefit of compacting the SIMD prefix (the EVEX prefix only requires 2 bits, rather than a byte to represent the SIMD prefix). In one embodiment, to support legacy SSE instructions using SIMD prefixes (66H, F2H, F3H) in both the legacy and EVEX prefix formats, these legacy SIMD prefixes are encoded in the SIMD prefix encoding field; and at runtime, they are expanded into legacy SIMD prefixes before being provided to the encoder's PLA (thus, the PLA can execute both legacy and EVEX formats of these legacy instructions without modification). While newer instructions can directly use the contents of the EVEX prefix encoding field as an opcode extension, some embodiments extend in a similar manner for consistency, but allow these legacy SIMD prefixes to specify different meanings. Alternative embodiments may redesign the PLA to support 2-bit SIMD prefix encodings, thus eliminating the need for expansion.

[0160] Alpha field 652 (EVEX byte 3, bit [7] - EH; also known as EVEX.EH, EVEX.rs, EVEX.RL, EVEX.WriteMaskControl, and EVEX.N; also illustrated as α) - As previously described, this field is context-dependent.

[0161] Beta field 654 (EVEX byte 3, bits [6:4] – SSS; also known as EVEX.s 2-0EVEX.r 2-0 , EVEX.rr1, EVEX.LL0, EVEX.LLB; also illustrated as βββ) – As mentioned earlier, this field is context-dependent.

[0162] REX' 710B - This is the remainder of the REX' field 710 and is the EVEX.V' bit field (EVEX byte 3, bit [3] - V') that can be used to encode the upper 16 or lower 16 of the extended 32-register set. This bit is stored in bit-reversed format. A value of 1 is used to encode the lower 16 registers. In other words, V'VVVV is formed by combining EVEX.V' and EVEX.vvvv.

[0163] Write mask field 670 (EVEX byte 3, bits [2:0] - kkk) - its contents specify the index of a register in the write mask register as previously described. In some embodiments, a special value of EVEX.kkk = 000 has special behavior that implies that no write mask is used for a particular instruction (this can be achieved in a variety of ways, including using a write mask that is hardwired to all ones or hardware that bypasses the masking hardware).

[0164] The real opcode field 730 (byte 4) is also called the opcode byte. A portion of the opcode is specified in this field.

[0165] The MOD R / M field 740 (byte 5) includes a MOD field 742, a Reg field 744, and an R / M field 746. As previously described, the contents of the MOD field 742 distinguish between memory access and non-memory access operations. The role of the Reg field 744 can be summarized into two cases: encoding a target register operand or a source register operand, or being treated as an opcode extension and not used to encode any instruction operand. The role of the R / M field 746 can include the following: encoding an instruction operand that references a memory address, or encoding a target register operand or a source register operand.

[0166] Scale, Index, Base (SIB) Byte (Byte 6) - As previously described, the contents of the scale field 650 are used for memory address generation. SIB.xxx 754 and SIB.bbb 756 - The contents of these fields have been referenced previously for register indices Xxxx and Bbbb.

[0167] Displacement field 662A (bytes 7-10) - When the MOD field 742 contains 10, bytes 7-10 are the displacement field 662A and it works the same as a traditional 32-bit displacement (disp32) and works on a byte granularity.

[0168] Displacement Factor Field 662B (Byte 7) - When the MOD field 742 contains 01, byte 7 is the displacement factor field 662B. The location of this field is the same as for the traditional x86 instruction set 8-bit displacement (disp8), which operates at a byte granularity. Because disp8 is sign-extended, it can only address offsets between -128 and 127 bytes. For a 64-byte cache line, disp8 uses 8 bits, which can be set to only four truly useful values: 128, -64, 0, and 64. Because a larger range is often needed, disp32 is used; however, disp32 requires 4 bytes. Unlike disp8 and disp32, the displacement factor field 662B is a reinterpretation of disp8. When the displacement factor field 662B is used, the actual displacement is determined by multiplying the contents of the displacement factor field by the size (N) of the memory operand being accessed. This type of displacement is called disp8*N. This reduces the average instruction length (a single byte is used for the displacement, but with a much larger range). This compressed displacement is based on the assumption that the effective displacement is a multiple of the granularity of the memory access, and therefore, the redundant low-order bits of the address offset do not need to be encoded. In other words, the displacement factor field 662B replaces the traditional x86 instruction set 8-bit displacement. Thus, the displacement factor field 662B is encoded in the same manner as the x86 instruction set 8-bit displacement (so there is no change in the ModRM / SIB encoding rules), with the only exception that disp8 is overloaded to disp8*N. In other words, there is no change in the encoding rules or encoding length, only in the hardware's interpretation of the displacement value (the hardware needs to scale the displacement by the size of the memory operand to obtain a byte address offset). The immediate field 672 operates as described above.

[0169] Full opcode field

[0170] Figure 7B is a block diagram illustrating the fields of a particular vector friendly instruction format 700 that make up the full opcode field 674, according to some embodiments. Specifically, the full opcode field 674 includes a format field 640, a base operation field 642, and a data element width (W) field 664. The base operation field 642 includes a prefix encoding field 725, an opcode map field 715, and a real opcode field 730.

[0171] Register index field

[0172] Figure 7Cis a block diagram illustrating fields of a particular vector friendly instruction format 700 that make up the register index field 644, according to some embodiments. Specifically, the register index field 644 includes a REX field 705, a REX' field 710, a MODR / M.reg field 744, a MODR / Mr / m field 746, a VVVV field 720, a xxx field 754, and a bbb field 756.

[0173] Enhanced Action Field

[0174] Figure 7D 6 is a block diagram illustrating the fields of a particular vector friendly instruction format 700 that make up the enhanced operation field 650, according to some embodiments. When the class (U) field 668 contains 0, it indicates EVEX.U0 (class A 668A); when it contains 1, it indicates EVEX.U1 (class B 668B). When U=0 and the MOD field 742 contains 11 (indicating a no memory access operation), the alpha field 652 (EVEX byte 3, bit [7] - EH) is interpreted as the rs field 652A. When the rs field 652A contains 1 (round 652A.1), the beta field 654 (EVEX byte 3, bits [6:4] - SSS) is interpreted as the round control field 654A. The round control field 654A includes a one-bit SAE field 656 and a two-bit round operation field 658. When the rs field 652A contains 0 (data transformation 652A.2), the beta field 654 (EVEX byte 3, bits [6:4]-SSS) is interpreted as a three-bit data transformation field 654B. When U=0 and the MOD field 742 contains 00, 01, or 10 (indicating a memory access operation), the alpha field 652 (EVEX byte 3, bits [7]-EH) is interpreted as an eviction hint (EH) field 652B and the beta field 654 (EVEX byte 3, bits [6:4]-SSS) is interpreted as a three-bit data manipulation field 654C.

[0175] When U=1, the alpha field 652 (EVEX byte 3, bit [7]-EH) is interpreted as the write mask control (Z) field 652C. When U=1 and the MOD field 742 contains 11 (indicating a no memory access operation), a portion of the beta field 654 (EVEX byte 3, bit [4]-S0) is interpreted as the RL field 657A; when it contains 1 (round 675A.1), the remainder of the beta field 654 (EVEX byte 3, bit [6-5]-S0) is interpreted as the RL field 657A. 2-1 ) is interpreted as the round operation field 659A, and when the RL field 657A contains 0 (VSIZE 657.A2), the remainder of the beta field 654 (EVEX byte 3, bits [6-5]-S 2-1) is interpreted as the vector length field 659B (EVEX byte 3, bits [6-5] - L 1-0 ). When U=1 and the MOD field 742 contains 00, 01, or 10 (indicating a memory access operation), the beta field 654 (EVEX byte 3, bits [6:4] - SSS) is interpreted as the vector length field 659B (EVEX byte 3, bits [6-5] - L 1-0 ) and the broadcast field 657B (EVEX byte 3, bit [4]-B).

[0176] Exemplary register architecture

[0177] Figure 8 is a block diagram of a register architecture 800 according to some embodiments. In the illustrated embodiment, there are 32 512-bit wide vector registers 810; these registers are referred to as zmm0 through zmm31. The low-order 256 bits of the lower 16 zmm registers are overlaid on registers ymm0-16. The low-order 128 bits of the lower 16 zmm registers (the low-order 128 bits of the ymm registers) are overlaid on registers xmm0-15. The specific vector-friendly instruction format 700 operates on these overlaid register files as shown in the following table.

[0178]

[0179] In other words, the vector length field 659B selects between a maximum length and one or more other shorter lengths, where each such shorter length is half the previous length; and instruction templates without a vector length field 659B operate on the maximum vector length. Additionally, in one embodiment, the class B instruction templates of the specific vector friendly instruction format 700 operate on packed or scalar single / double precision floating point data and packed or scalar integer data. Scalar operations are operations performed on the lowest-order data element position in an zmm / ymm / xmm register; higher-order data element positions are either left the same as they were before the instruction or are zeroed, depending on the embodiment.

[0180] Write mask registers 815 - In the illustrated embodiment, there are eight write mask registers (k0 through k7), each 64 bits in size. In an alternative embodiment, the write mask registers 815 are 16 bits in size. As previously mentioned, in some embodiments, vector mask register k0 can be used as a write mask; while the encoding that would normally indicate k0 is used for a write mask, it selects a hardwired write mask of 0xffff, effectively disabling write masking for that instruction.

[0181] General purpose registers 825 - In the illustrated embodiment, there are sixteen 64-bit general purpose registers that are used with existing x86 addressing modes to address memory operands. These registers are referenced by the names RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 through R15.

[0182] Scalar floating point stack register file (x87 stack) 845, on which is aliased the MMX packed integer flat register file 850 - in the illustrated embodiment, the x87 stack is an eight-element stack used to perform scalar floating point operations on 32 / 64 / 80-bit floating point data using the x87 instruction set extension; while the MMX registers are used to perform operations on 64-bit packed integer data, as well as to hold operands for some operations performed between MMX and XMM registers.

[0183] Alternative embodiments may use wider or narrower registers. Additionally, alternative embodiments may use more, fewer, or different register files and registers.

[0184] Exemplary Core Architectures, Processors, and Computer Architectures

[0185] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, implementations of such cores may include: 1) a general-purpose in-order core intended for general-purpose computing; 2) a high-performance general-purpose out-of-order core intended for general-purpose computing; 3) a specialized core intended primarily for graphics and / or scientific (throughput) computing. Different processor implementations may include: 1) a CPU including one or more general-purpose in-order cores intended for general-purpose computing and / or one or more general-purpose out-of-order cores intended for general-purpose computing; and 2) a coprocessor including one or more specialized cores intended primarily for graphics and / or scientific (throughput) computing. Such different processors lead to different computer system architectures, which may include: 1) the coprocessor is on a separate chip from the CPU; 2) the coprocessor is in the same package as the CPU, on a separate die; 3) the coprocessor is on the same die as the CPU (in which case such coprocessor is sometimes referred to as dedicated logic, such as integrated graphics and / or scientific (throughput) logic, or as a dedicated core); and 4) a system on a chip, which may include the described CPU (sometimes referred to as application core(s) or application processor(s), the coprocessor described above, and additional functionality on the same die. An exemplary core architecture is described next, followed by a description of an exemplary processor and computer architecture.

[0186] Exemplary core architecture

[0187] In-order and out-of-order core block diagram

[0188] Figure 9A is a block diagram illustrating both an exemplary in-order pipeline and an exemplary register renaming, out-of-order issue / execution pipeline, according to some embodiments of the invention. Figure 9B is a block diagram illustrating both an exemplary embodiment of an in-order architecture core and an exemplary register renaming, out-of-order issue / execution architecture core to be included in a processor according to some embodiments of the invention. Figures 9A-9B The solid boxes in illustrate the in-order pipeline and in-order core, while the optional addition of dashed boxes illustrates the register renaming, out-of-order issue / execution pipeline, and core. Considering that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.

[0189] exist Figure 9A , the processor pipeline 900 includes a fetch stage 902, a length decode stage 904, a decode stage 906, an allocation stage 908, a rename stage 910, a schedule (also called dispatch or issue) stage 912, a register read / memory read stage 914, an execute stage 916, a write back / memory write stage 918, an exception handling stage 922, and a commit stage 924.

[0190] Figure 9B Processor core 990 is shown to include a front end unit 930 coupled to an execution engine unit 950, and both are coupled to a memory unit 970. Core 990 can be a reduced instruction set computing (RISC) core, a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, core 990 can be a specialized core, such as a network or communication core, a compression engine, a coprocessor core, a general purpose computing graphics processing unit (GPGPU) core, a graphics core, or the like.

[0191] The front end unit 930 includes a branch prediction unit 932, which is coupled to an instruction cache unit 934, which is coupled to an instruction translation lookaside buffer (TLB) 936, which is coupled to an instruction fetch unit 938, which is coupled to a decode unit 940. The decode unit 940 (or decoder) can decode the instruction and generate one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals as output. These micro-operations, microcode entry points, microinstructions, other instructions, or other control signals are decoded from the original instruction, or otherwise reflect the original instruction, or are derived from the original instruction. The decode unit 940 can be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLA), microcode read-only memory (ROM), and the like. In one embodiment, core 990 includes a microcode ROM or other medium that stores microcode for certain macroinstructions (e.g., in decode unit 940 or otherwise within front end unit 930). Decode unit 940 is coupled to rename / allocator unit 952 in execution engine unit 950.

[0192] The execution engine unit 950 includes a rename / allocator unit 952 coupled to a retirement unit 954 and a set of one or more scheduler units 956. The scheduler unit(s) 956 represent any number of different schedulers, including reservation stations, central instruction windows, and the like. The scheduler unit(s) 956 are coupled to the physical register file(s) 958. Each of the physical register file units 958 represents one or more physical register files, different ones of which store one or more different data types, such as scalar integers, scalar floating point, packed integers, packed floating point, vector integers, vector floating point, status (e.g., an instruction pointer that is the address of the next instruction to be executed), and the like. In one embodiment, the physical register file units 958 include a vector register unit, a write mask register unit, and a scalar register unit. These register units may provide architectural vector registers, vector mask registers, and general purpose registers. The physical register file(s) unit(s) 958 are overlaid with the retirement unit 954 to illustrate various ways in which register renaming and out-of-order execution may be implemented (e.g., using reorder buffer(s) and retirement register file(s); using future file(s), history buffer(s), and retirement register file(s); using register maps and pools of registers; etc.). The retirement unit 954 and the physical register file(s) unit(s) 958 are coupled to the execution cluster(s) 960. The execution cluster(s) 960 include a set of one or more execution units 962 and a set of one or more memory access units 964. The execution units 962 can perform various operations (e.g., shifts, additions, subtractions, multiplications) on various types of data (e.g., scalar floating point, packed integer, packed floating point, vector integer, vector floating point). While some embodiments may include several execution units dedicated to a particular function or set of functions, other embodiments may include only one execution unit or multiple execution units that all perform all functions. Scheduler unit(s) 956, physical register file(s) 958, and execution cluster(s) 960 are shown as potentially multiple because some embodiments create separate pipelines for certain types of data / operations (e.g., a scalar integer pipeline, a scalar floating point / packed integer / packed floating point / vector integer / vector floating point pipeline, and / or a memory access pipeline, each with its own scheduler unit, physical register file unit, and / or execution cluster—and in the case of a separate memory access pipeline, some embodiments are implemented where only the execution cluster for that pipeline has memory access unit(s) 964). It should also be understood that where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution, with the rest being in-order.

[0193] A set of memory access units 964 is coupled to a memory unit 970, which includes a data TLB unit 972, which is coupled to a data cache unit 974, which is coupled to a level 2 (L2) cache unit 976. In one exemplary embodiment, the memory access units 964 may include a load unit, a store address unit, and a store data unit, each of which is coupled to the data TLB unit 972 in the memory unit 970. The instruction cache unit 934 is further coupled to a level 2 (L2) cache unit 976 in the memory unit 970. The L2 cache unit 976 is coupled to one or more other levels of cache and ultimately to main memory.

[0194] As an example, the exemplary register renaming, out-of-order issue / execution core architecture may implement pipeline 900 as follows: 1) instruction fetch 938 performs fetch and length decode stages 902 and 904; 2) decode unit 940 performs decode stage 906; 3) rename / allocator unit 952 performs allocate stage 908 and rename stage 910; 4) (one or more) scheduler units 956 perform schedule stage 912; 5) (one or more) physical register file units 958 and memory units 970 perform register read / memory read stage 914; execution cluster 960 performs execute stage 916; 6) memory units 970 and (one or more) physical register file units 958 perform write back / memory write stage 918; 7) various units may be involved in exception handling stage 922; and 8) retirement unit 954 and (one or more) physical register file units 958 perform commit stage 924.

[0195] Core 990 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions that have been added with later versions); the MIPS instruction set from MIPS Technologies, Inc. of Sunnyvale, California; the ARM instruction set from ARM Holdings, Inc. of Sunnyvale, California (with optional additional extensions, such as NEON)), including the instruction(s) described herein. In one embodiment, core 990 includes logic to support packed data instruction set extensions (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using packed data.

[0196] It should be understood that a core can support multithreading (executing two or more parallel sets of operations or threads) and can support multithreading in a variety of ways, including time-sliced ​​multithreading, simultaneous multithreading (where a single physical core provides a logical core for each thread that the physical core is multithreading at the same time), or a combination thereof (e.g., time-sliced ​​fetch and decode followed by simultaneous multithreading, such as Hyperthreading technology).

[0197] Although register renaming is described in the context of out-of-order execution, it should be understood that register renaming can be used in an in-order architecture. Although the illustrated embodiment of the processor also includes separate instruction and data cache units 934 / 974 and a shared L2 cache unit 976, alternative embodiments may have a single internal cache for both instructions and data, such as a level 1 (L1) internal cache, or multiple levels of internal cache. In some embodiments, the system may include a combination of internal caches and external caches external to the core and / or processor. Alternatively, all caches may be external to the core and / or processor.

[0198] Specific exemplary in-order core architecture

[0199] Figures 10A-10B The figure shows a block diagram of a more specific exemplary in-order core architecture, which would be one of several logic blocks in a chip (including other cores of the same and / or different types). The logic block communicates with some fixed-function logic, memory I / O interfaces, and other necessary I / O logic via a high-bandwidth interconnect network (e.g., a ring network), depending on the application.

[0200] Figure 10A 10 is a block diagram of a single processor core and its connections to an on-chip interconnect network 1002 and to its local subset 1004 of a level 2 (L2) cache according to some embodiments of the present invention. In one embodiment, the instruction decoder 1000 supports the x86 instruction set with the packed data instruction set extension. The L1 cache 1006 allows low-latency access to cache memory in the scalar and vector units. While in one embodiment (to simplify the design), the scalar unit 1008 and the vector unit 1010 use separate register sets (scalar registers 1012 and vector registers 1014, respectively) and data transferred between them is written to memory and then read back from the level 1 (L1) cache 1006, alternative embodiments of the present invention may use different schemes (e.g., using a single register set or including a communication path that allows data to be transferred between the two register files without being written and read back).

[0201] The local subset 1004 of the L2 cache is part of the global L2 cache, which is divided into separate local subsets, one for each processor core. Each processor core has a direct access path to its own local subset 1004 of the L2 cache. Data read by a processor core is stored in its L2 cache subset 1004 and can be accessed quickly, in parallel with other processor cores accessing their own local L2 cache subsets. Data written by a processor core is stored in its own L2 cache subset 1004 and flushed from other subsets when necessary. The ring network ensures the consistency of shared data. The ring network is bidirectional to allow agents such as processor cores, L2 caches, and other logic blocks to communicate with each other within the chip. Each ring data path is 1012 bits wide in each direction.

[0202] Figure 10B According to some embodiments of the present invention Figure 10A An expanded view of a portion of a processor core. Figure 10B Includes the L1 data cache 1006A portion of the L1 cache 1004, as well as more details about the vector unit 1010 and vector registers 1014. Specifically, the vector unit 1010 is a 16-wide vector processing unit (VPU) (see 16-wide ALU 1028) that executes one or more of integer, single-precision floating-point, and double-precision floating-point instructions. The VPU supports swabbing register inputs using the swabbing unit 1020, numerical conversion using the numerical conversion units 1022A-B, and copying of memory inputs using the copy unit 1024. Write mask register 1026 allows assertion of result vector writes.

[0203] Figure 11 is a block diagram of a processor 1100 that may have more than one core, may have an integrated memory controller, and may have integrated graphics, according to some embodiments of the present invention. Figure 11 The solid line box in the figure illustrates a processor 1100 having a single core 1102A, a system agent 1110, and a set of one or more bus controller units 1116, while the optional addition of the dashed line box illustrates an alternative processor 1100 having multiple cores 1102A-N, a set of one or more integrated memory control units 1114 in the system agent unit 1110, and dedicated logic 1108.

[0204] Thus, different implementations of processor 1100 may include: 1) a CPU in which specialized logic 1108 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores) and cores 1102A-N are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of both); 2) a coprocessor in which cores 1102A-N are a large number of specialized cores primarily intended for graphics and / or scientific (throughput); and 3) a coprocessor in which cores 1102A-N are a large number of general-purpose in-order cores. Thus, processor 1100 may be a general-purpose processor, a coprocessor, or a specialized processor, such as a network or communications processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput many-integrated-core (MIC) coprocessor (including 30 or more cores), an embedded processor, and the like. The processor may be implemented on one or more chips. Processor 1100 may be part of and / or implemented on one or more substrates using any of several process technologies, such as BiCMOS, CMOS, or NMOS.

[0205] The memory hierarchy may include one or more levels of cache within the core, a set or one or more shared cache units 1106, and external memory (not shown) coupled to the set of integrated memory controller units 1114. The set of shared cache units 1106 may include (one or more) intermediate level caches, such as level 2 (L2), level 3 (L3), level 4 (4), or other levels of cache, a last level cache (LLC), and / or combinations thereof. While in one embodiment a ring-based interconnect unit 1112 interconnects the integrated graphics logic 1108 (which is an example of specialized logic and is also referred to herein as specialized logic), the set of shared cache units 1106, and the system agent unit 1110 / (one or more) integrated memory controller units 1114, alternative embodiments may use any number of well-known techniques to interconnect such units. In one embodiment, coherency is maintained between the one or more cache units 1106 and the cores 1102A-N.

[0206] In some embodiments, one or more of the cores 1102A-N may be capable of multi-threaded processing. System agent 1110 includes components that coordinate and operate cores 1102A-N. System agent unit 1110 may include, for example, a power control unit (PCU) and a display unit. The PCU may be or may include logic and components required to regulate the power state of cores 1102A-N and integrated graphics logic 1108. The display unit is used to drive one or more externally connected displays.

[0207] The cores 1102A-N may be homogeneous or heterogeneous with respect to architectural instruction sets; that is, two or more of the cores 1102A-N may be capable of executing the same instruction set, while others may be capable of executing only a subset of the instruction set or a different instruction set.

[0208] Exemplary Computer Architecture

[0209] Figure 12-15 is a block diagram of an exemplary computer architecture. Other system designs and configurations known in the art for laptop computers, desktop computers, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and various other electronic devices are also suitable. In short, many types of systems or electronic devices that can include the processors and / or other execution logic disclosed herein are generally suitable.

[0210] Now refer to Figure 12, which shows a block diagram of a system 1200 according to one embodiment of the present invention. System 1200 includes one or more processors 1210, 1215 coupled to a controller hub 1220. In one embodiment, controller hub 1220 includes a graphics memory controller hub (GMCH) 1290 and an input / output hub (IOH) 1250 (which may be on separate chips); GMCH 1290 includes memory and a graphics processor coupled to memory 1240 and a coprocessor 1245; and IOH 1250 couples input / output (I / O) devices 1260 to GMCH 1290. Alternatively, one or both of the memory and graphics controller are integrated within the processor (as described herein), memory 1240 and coprocessor 1245 are directly coupled to processor 1210, and controller hub 1220 and IOH 1250 are on a single chip.

[0211] The optional additional processor 1215 is Figure 12 Each processor 1210 , 1215 may include one or more of the processing cores described herein and may be some version of processor 1100 .

[0212] The memory 1240 may be, for example, dynamic random-access memory (DRAM), phase change memory (PCM), or a combination thereof. For at least one embodiment, the controller hub 1220 communicates with the controller(s) 1210 and 1215 via a multi-drop bus (e.g., a front-side bus (FSB)), a point-to-point interface (e.g., a QuickPath Interconnect (QPI)), or similar connection 1295.

[0213] In one embodiment, the coprocessor 1245 is a special purpose processor, such as a high throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, etc. In one embodiment, the controller hub 1220 may include an integrated graphics accelerator.

[0214] There may be a variety of differences between the physical resources 1210, 1215 in terms of a range of value metrics including architectural characteristics, microarchitectural characteristics, thermal characteristics, power consumption characteristics, and the like.

[0215] In one embodiment, processor 1210 executes instructions that control general types of data processing operations. Embedded within these instructions may be coprocessor instructions. Processor 1210 recognizes these coprocessor instructions as being of a type that should be executed by an attached coprocessor 1245. Accordingly, processor 1210 issues these coprocessor instructions (or control signals representing coprocessor instructions) to coprocessor 1245 over a coprocessor bus or other interconnect. Coprocessor(s) 1245 accept and execute the received coprocessor instructions.

[0216] Now refer to Figure 13 , which shows a block diagram of a first more specific exemplary system 1300 according to an embodiment of the present invention. Figure 13 As shown in FIG, multiprocessor system 1300 is a point-to-point interconnect system and includes a first processor 1370 and a second processor 1380 coupled via a point-to-point interconnect 1350. Each of processors 1370 and 1380 may be some version of processor 1100. In some embodiments, processors 1370 and 1380 are processors 1210 and 1215, respectively, and coprocessor 1338 is coprocessor 1245. In another embodiment, processors 1370 and 1380 are processor 1210 and coprocessor 1245, respectively.

[0217] Processors 1370 and 1380 are shown as including integrated memory controller (IMC) units 1372 and 1382, respectively. Processor 1370 also includes point-to-point (PP) interfaces 1376 and 1378 as part of its bus controller unit; similarly, second processor 1380 includes PP interfaces 1386 and 1388. Processors 1370, 1380 can exchange information via point-to-point (PP) interface 1350 using PP interface circuits 1378, 1388. Figure 13 As shown in FIG, IMCs 1372 and 1382 couple the processors to respective memories, namely, memory 1332 and memory 1334, which may be portions of main memory locally attached to the respective processors.

[0218] Processors 1370, 1380 may each exchange information with a chipset 1390 via individual PP interfaces 1352, 1354 using point-to-point interface circuits 1376, 1394, 1386, 1398. Chipset 1390 may optionally exchange information with a coprocessor 1338 via a high-performance interface 1392. In one embodiment, coprocessor 1338 is a special-purpose processor, such as a high-throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, or the like.

[0219] A shared cache (not shown) may be included in either processor, or external to both processors but connected to the processors via the PP interconnect, such that local cache information of either or both processors may be stored in the shared cache when the processors are placed in a low power mode.

[0220] Chipset 1390 may be coupled to first bus 1316 via interface 1396. In one embodiment, first bus 1316 may be a Peripheral Component Interconnect (PCI) bus, or a bus such as a PCI Express bus or another third generation I / O interconnect bus, although the scope of the invention is not limited in this respect.

[0221] like Figure 13 As shown in , various I / O devices 1314 may be coupled to the first bus 1316, and a bus bridge 1318 coupling the first bus 1316 to a second bus 1320. In one embodiment, one or more additional processors 1315, such as coprocessors, high throughput MIC processors, GPGPUs, accelerators (e.g., graphics accelerators or digital signal processing (DSP) units), field programmable gate arrays, or any other processors, are coupled to the first bus 1316. In one embodiment, the second bus 1320 may be a low pin count (LPC) bus. Various devices may be coupled to the second bus 1320, including, for example, a keyboard and / or mouse 1322, communication devices 1327, and a storage unit 1328, such as a disk drive or other mass storage device, which in one embodiment may include instructions / code and data 1330. Additionally, an audio I / O 1324 may be coupled to the second bus 1320. Note that other architectures are possible. For example, instead of Figure 13 Instead of a point-to-point architecture, the system can implement a multi-point branch bus or other such architecture.

[0222] Now refer to Figure 14 , which shows a block diagram of a second more specific exemplary system 1400 according to an embodiment of the present invention. Figure 13 and Figure 14 Like elements in the text are numbered like, and Figure 13 Some aspects of Figure 14 Omitted to avoid ambiguity Figure 14 other aspects.

[0223] Figure 14The diagram illustrates that processors 1370, 1380 may respectively include integrated memory and I / O control logic ("CL") 1472 and 1482. Thus, CL 1472, 1482 includes an integrated memory controller unit and includes I / O control logic. Figure 14 The diagram shows that not only are memories 1332, 1334 coupled to CLs 1472, 1482, but I / O devices 1414 are also coupled to control logic 1472, 1482. Legacy I / O devices 1415 are coupled to chipset 1390.

[0224] Now refer to Figure 15 , which shows a block diagram of a SoC 1500 according to an embodiment of the present invention. Figure 11 Similar elements in the figure are numbered similarly. In addition, dashed boxes are optional features on more advanced SoCs. Figure 15 In the embodiment, interconnect unit(s) 1502 couples to: application processor(s) 1510, which includes a set of one or more cores 1102A-N, and shared cache unit(s) 1106, wherein cores 1102A-N include cache units 1104A-N; system agent unit 1110; bus controller unit(s) 1116; integrated memory controller unit(s) 1114; a set or one or more coprocessors 1520, which may include integrated graphics logic, image processors, audio processors, and video processors; static random access memory (SRAM) unit 1530; direct memory access (DMA) unit 1532; and display unit 1540 for coupling to one or more external displays. In one embodiment, coprocessor(s) 1520 include specialized processors, such as network or communication processors, compression engines, GPGPUs, high-throughput MIC processors, embedded processors, and the like.

[0225] Embodiments of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of such implementations. Embodiments of the present invention may be implemented as a computer program or program code executed on a programmable system comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0226] Program code, such as Figure 13The code 1330 shown in can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor, such as a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0227] Program code can be implemented in high-level procedural or object-oriented programming languages ​​to communicate with the processing system. If desired, program code can also be implemented in assembly or machine language. In fact, the mechanism described herein is not limited to any specific programming language in scope. In any case, the language can be a compiled or interpreted language.

[0228] One or more aspects of at least one embodiment may be implemented as representative instructions stored on a machine-readable medium representing various logic within a processor, which, when read by a machine, causes the machine to fabricate the logic to perform the techniques described herein. Such representations, known as "IP cores," may be stored on a tangible, machine-readable medium and supplied to various customers or manufacturing facilities to load into the fabrication machines that actually make the logic or processor.

[0229] Such machine-readable storage media may include, but are not limited to, a non-transitory tangible arrangement of an article manufactured or formed by a machine or apparatus, including a storage medium such as a hard disk, any other type of disk including a floppy disk, an optical disk, a compact disk read-only memory (CD-ROM), a compact disk rewritable (CD-RW), and a magneto-optical disk, a semiconductor device such as a read-only memory (ROM), a random access memory (RAM), such as a dynamic random access memory (DRAM), a static random access memory (SRAM), an erasable programmable read-only memory (EEPROM), a flash memory, an electrically erasable programmable read-only memory (EEPROM), a phase change memory (PCM), a magnetic or optical card, or any other type of medium suitable for storing electronic instructions.

[0230] Therefore, embodiments of the present invention also include non-transitory tangible machine-readable media containing instructions or design data, such as hardware description language (HDL), that define the features of the structures, circuits, devices, processors, and / or systems described herein. Such embodiments may also be referred to as program products.

[0231] Emulation (including binary conversion, code deformation, etc.)

[0232] In some cases, an instruction converter can be used to convert instructions from a source instruction set into a target instruction set. For example, the instruction converter can convert instructions (e.g., using static binary conversion, dynamic binary conversion including dynamic compilation), deform, simulate or otherwise convert instructions into one or more other instructions to be processed by the core. The instruction converter can be implemented in software, hardware, firmware or a combination thereof. The instruction converter can be on the processor, outside the processor or partly on the processor and partly outside the processor.

[0233] Figure 161 is a block diagram illustrating a method for converting binary instructions in a source instruction set into binary instructions in a target instruction set using a software instruction converter according to some embodiments of the present invention. In the illustrated embodiment, the instruction converter is a software instruction converter, although alternatively, the instruction converter can be implemented in software, firmware, hardware, or various combinations thereof. Figure 16 It is shown that a program in a high-level language 1602 can be compiled using an x86 compiler 1604 to generate x86 binary code 1606 that can be natively executed by a processor having at least one x86 instruction set core 1616. A processor having at least one x86 instruction set core 1616 represents any processor that can perform substantially the same functions as an Intel processor having at least one x86 instruction set core by compatibly executing or otherwise processing (1) a substantial portion of the instruction set of an Intel x86 instruction set core or (2) an object code version of an application or other software running on an Intel processor having at least one x86 instruction set core to achieve substantially the same results as an Intel processor having at least one x86 instruction set core. The x86 compiler 1604 represents a compiler that can operate to generate x86 binary code 1606 (e.g., object code) that can be executed on a processor having at least one x86 instruction set core 1616 with or without additional linking processing. Similarly, Figure 16 A program in a high-level language 1602 is shown to be compiled using an alternative instruction set compiler 1608 to generate alternative instruction set binary code 1610, which is natively executable by a processor 1614 that does not have at least one x86 instruction set core (e.g., a processor having a core that executes the MIPS instruction set of MIPS Technologies, Inc. of Sunnyvale, California and / or the ARM instruction set of ARM Holdings, Inc. of Sunnyvale, California). An instruction converter 1612 is used to convert x86 binary code 1606 into code that is natively executable by a processor 1614 that does not have an x86 instruction set core. This converted code is unlikely to be identical to the alternative instruction set binary code 1610, as an instruction converter capable of doing so is difficult to create; however, the converted code will implement general operations and be composed of instructions from the alternative instruction set. Thus, instruction converter 1612 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device that does not have an x86 instruction set processor or core to execute x86 binary code 1606 through emulation, simulation, or any other process.

[0234] Further examples

[0235] Example 1 describes an exemplary processor, comprising: an extraction circuit for extracting an instruction, the instruction having a field specifying an opcode and a location of a first source vector including N single-precision elements and a destination vector including at least N 16-bit floating-point elements, the opcode instructing an execution circuit to convert each element of the specified source vector to a 16-bit floating-point format and store each converted element in a corresponding location of the specified destination vector, the conversion including truncation and rounding as necessary; a decoding circuit for decoding the extracted instruction; and an execution circuit for responding to the instruction according to the opcode.

[0236] Example 2 includes the essence of the exemplary processor as described in Example 1, wherein: the instruction further specifies the location of a second source vector including N single-precision elements; the specified destination vector includes 2 times N 16-bit floating-point elements, the first half and the second half of which correspond to the first source vector and the second source vector, respectively; and the opcode instructs the processor to convert each element of the specified first and second source vectors to 16-bit floating-point format and store each converted element in a corresponding position of the specified destination vector, the conversion including truncation and rounding as necessary.

[0237] Example 3 includes the essence of the exemplary processor of Example 1, wherein the location of each specified source vector and destination vector is in a register or in a memory.

[0238] Example 4 includes the substance of the exemplary processor of Example 1, wherein the 16-bit floating point format includes a sign bit, an 8-bit exponent, and a mantissa, the mantissa including 7 explicit bits and an eighth implicit bit.

[0239] Example 5 includes the substance of the exemplary processor of Example 1, wherein N is specified by the instruction and has a value of one of 4, 8, 16, and 32.

[0240] Example 6 includes the substance of the exemplary processor of Example 1, wherein when the execution circuitry performs rounding, it performs according to a round-to-nearest-even rule.

[0241] Example 7 includes the substance of the exemplary processor of Example 1, wherein the 16-bit floating point format is bfloat16 or binary16.

[0242] Example 8 includes the substance of the exemplary processor of Example 1, wherein the execution circuitry generates all N elements of the specified target in parallel.

[0243] Example 9 describes an exemplary method performed by a processor, the method comprising: using an extraction circuit to extract an instruction, the instruction having a field specifying an opcode and a location of a first source vector including N single-precision elements and a destination vector including at least N 16-bit floating-point elements, the opcode instructing an execution circuit to convert each element of the specified source vector to a 16-bit floating-point format and store each converted element in a corresponding location of the specified destination vector, the conversion including truncation and rounding as necessary; using a decoding circuit to decode the extracted instruction; and using the execution circuit to respond to the instruction according to the opcode.

[0244] Example 10 includes the essence of the exemplary method as described in Example 9, wherein: the instruction further specifies the location of a second source vector including N single-precision elements; the specified destination vector includes 2 times N 16-bit floating-point elements, the first half and the second half of which correspond to the first source vector and the second source vector, respectively; and the opcode instructs the execution circuit to convert each element of the specified first and second source vectors into 16-bit floating-point format, and store each converted element into a corresponding position of the specified destination vector, the conversion including truncation and rounding as necessary.

[0245] Example 11 includes the essence of the exemplary method of Example 9, wherein the location of each specified source vector and destination vector is in a register or in a memory.

[0246] Example 12 includes the essence of the exemplary method of Example 9, wherein the 16-bit floating point format includes a sign bit, an 8-bit exponent, and a mantissa, the mantissa including 7 explicit bits and an eighth implicit bit.

[0247] Example 13 includes the substance of the exemplary method of Example 9, wherein N is specified by the instruction and has a value of one of 4, 8, 16, and 32.

[0248] Example 14 includes the essence of the exemplary method of Example 9, wherein when the execution circuit performs rounding, it performs according to the rounding rule promulgated as IEEE 754 round to nearest even.

[0249] Example 15 includes the essence of the exemplary method of Example 9, wherein the 16-bit floating point format is bfloat16 or binary16.

[0250] Example 16 includes the substance of the exemplary method of Example 9, wherein the execution circuit generates all N elements of the specified target in parallel.

[0251] Example 17 describes an exemplary non-transitory machine-readable medium containing instructions that, when executed by a processor, cause the processor to respond by: extracting, using extraction circuitry, an instruction having a field specifying an opcode and a location of a first source vector comprising N single-precision elements and a destination vector comprising at least N 16-bit floating-point elements, the opcode instructing execution circuitry to convert each element of the specified source vector to 16-bit floating-point format and store each converted element in a corresponding location of the specified destination vector, the conversion including truncation and rounding as necessary; decoding, using decoding circuitry, the extracted instruction; and responding to the instruction, using execution circuitry, according to the opcode.

[0252] Example 18 includes the substance of the exemplary non-transitory machine-readable medium as described in Example 17, wherein: the instruction further specifies the location of a second source vector including N single-precision elements; the specified destination vector includes 2 times N 16-bit floating-point elements, the first half and the second half of which correspond to the first source vector and the second source vector, respectively; and the opcode instructs the execution circuit to convert each element of the specified first and second source vectors to 16-bit floating-point format and store each converted element in a corresponding position of the specified destination vector, the conversion including truncation and rounding as necessary.

[0253] Example 19 includes the substance of the exemplary non-transitory machine-readable medium of Example 17, wherein the location of each specified source vector and destination vector is in a register or in a memory.

[0254] Example 20 includes the substance of the exemplary non-transitory machine-readable medium of Example 17, wherein when the execution circuitry performs rounding, it performs according to a round-to-nearest-even rule.

Claims

1. A processor, comprising: Control register, used to specify rounding mode; an extraction circuit for extracting a format conversion instruction; a decoding circuit configured to decode the format conversion instruction, the format conversion instruction having an opcode, a first field, a second field, and a third field, the first field being configured to specify a source vector register, the second field being configured to specify a mask, and the third field being configured to specify a destination vector register, the source vector register being configured to store a source vector having a plurality of 32-bit single-precision floating-point data elements, wherein the mask has a plurality of mask bits, and the plurality of mask bits respectively correspond to a plurality of data element positions of the destination vector register; an execution circuit coupled to the decoding circuit, configured to execute an operation corresponding to the format conversion instruction, the operation comprising: For the data element positions in the target vector register where the corresponding mask bit is 1, the following operation is performed: converting corresponding 32-bit single-precision floating-point data elements of the source vector into corresponding 16-bit floating-point data elements according to the rounding mode specified by the control register, the 16-bit floating-point data elements having a format including a sign bit, an 8-bit exponent, seven explicit mantissa bits, and one implicit mantissa bit; and storing the 16-bit floating point data element in a data element position of the destination vector register; and For a data element position in the target vector register whose corresponding mask bit is 0, the value in the data element position is not changed.

2. The processor according to claim 1, wherein: The target vector register has a plurality of data element positions that do not correspond to the plurality of mask bits.

3. The processor according to claim 2, wherein: The number of data element positions that do not correspond to the plurality of mask bits is the same as the number of data element positions that correspond to the plurality of mask bits.

4. The processor according to any one of claims 1 to 3, wherein: The mask is a first register in a register set having a second register, the processor not supporting use of the second register as a mask.

5. The processor according to any one of claims 1 to 4, wherein: The source vector is one of: a 128-bit source vector, a 256-bit source vector, a 512-bit source vector, or a 1024-bit source vector.

6. The processor according to any one of claims 1 to 5, wherein: The format conversion instruction allows the source vector to be either 128 bits or 512 bits.

7. The processor according to any one of claims 1 to 6, wherein: The format is bfloat16 format.

8. The processor according to any one of claims 1 to 7, wherein: The processor is a general purpose CPU core.

9. The processor according to any one of claims 1 to 8, wherein: The processor is a Reduced Instruction Set Computing (RISC) processor.

10. A method comprising: Specify the rounding mode in the control register; Extract format conversion instructions; decoding the format conversion instruction, the format conversion instruction having an opcode, a first field, a second field, and a third field, the first field being used to specify a source vector register, the second field being used to specify a mask, and the third field being used to specify a destination vector register, the source vector register being used to store a source vector having a plurality of 32-bit single-precision floating-point data elements, wherein the mask has a plurality of mask bits, the plurality of mask bits respectively corresponding to a plurality of data element positions of the destination vector register; Executing an operation corresponding to the format conversion instruction, the operation including: For the data element positions in the target vector register where the corresponding mask bit is 1, the following operation is performed: converting corresponding 32-bit single-precision floating-point data elements of the source vector into corresponding 16-bit floating-point data elements according to the rounding mode specified by the control register, the 16-bit floating-point data elements having a format including a sign bit, an 8-bit exponent, seven explicit mantissa bits, and one implicit mantissa bit; and storing the 16-bit floating point data element in a data element position of the destination vector register; and For a data element position in the target vector register whose corresponding mask bit is 0, the value in the data element position is not changed.

11. A computer program comprising instructions which, when executed by a computer, cause the computer to perform the method according to claim 10.

12. A computer-readable storage medium having stored thereon the computer program according to claim 11.

13. A method comprising: Write the first source vector sequentially into the entry of the destination vector; as well as Write the second source vector into the other entry of the destination vector; The sum of the width of the first source vector and the width of the second source vector is less than or equal to the width of the target vector.

14. The method according to claim 13, further comprising: When the sum of the width of the first source vector and the width of the second source vector is smaller than the width of the target vector, a truncation operation is performed on the target vector.

15. The method according to claim 13, further comprising: When the sum of the width of the first source vector and the width of the second source vector is smaller than the width of the target vector, a zeroing operation is performed on the target vector. 16 . A computer-readable storage medium having instructions stored thereon, which, when executed by a processor, cause the processor to perform the method according to claim 13 .

17. An apparatus comprising components for performing the method of any one of claims 13 to 15.

Citation Information

Cited By

  • Instruction execution method and device and processor

    CN120950126A

  • Instruction execution method and apparatus and processor

    CN120950126B