Processing device, method and computer program for vector compound instructions
Patent Information
- Application Number
- JP2024500529
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-07-21
- Filing Date
- 2022-06-22
- Publication Date
- 2025-06-23
AI Technical Summary
Existing processing technologies are inefficient in performing repeated operations on data elements due to the sequential processing of individual elements, which can significantly impact performance.
A processing device and method that utilizes specialized circuitry to respond to vector composition instructions, allowing simultaneous processing of multiple data elements from multiple source registers, including source and destination registers, through decoding and composition operations.
This approach enhances processing efficiency by enabling parallel or sequential performance of composition operations on multiple data elements, reducing the overhead of sequential processing and improving performance.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] A processing unit may include processing circuitry to perform vector processing operations. Such operations may involve operations on elements of vector registers using processing circuitry designed to enable efficient processing of one or more instructions.
[0002] According to some example configurations, a processing device is provided that includes a decoding circuit for decoding instructions and a processing circuit for selectively applying a vector processing operation specified by the instruction to an input data vector including a plurality of input data items at respective positions within the input data vector, wherein the decoding circuit is configured to, in response to a vector composition instruction specifying a plurality of source vector registers, each including source data elements at a plurality of data element positions, one or more further source vector registers, and one or more destination registers, generate control signals that cause the processing circuit to, for each data element position of the plurality of data element positions, extract a first source data element from the data element position of each source vector register and extract a second source data element from the one or more further source vector registers and perform a composition operation to generate a result data element, the result data element being calculated by combining each element of the first source data element and the second source data element, and storing the result data element in a data element position of the one or more destination registers.
[0003] According to another exemplary configuration, there is provided a method of operating a processing device comprising a decode circuit for decoding instructions and a processing circuit for selectively applying a vector processing operation specified by the instruction to an input data vector including a plurality of input data items at respective positions within the input data vector, the method including: using the decode circuit to generate control signals that, in response to a vector combine instruction specifying a plurality of source vector registers, each including a source data element at a plurality of data element positions, one or more further source vector registers, and one or more destination registers, cause the processing circuit to perform, for each data element position of the plurality of data element positions, the steps of extracting a first source data element from a data element position of each source vector register, extracting a second source data element from the one or more further source vector registers, performing a combine operation to generate a result data element, the result data element being computed by combining each element of the first source data element and the second source data element, and storing the result data element in a data element position of the one or more destination registers.
[0004] According to another exemplary arrangement, there is provided a computer program for controlling a host processing apparatus to provide an instruction execution environment, the computer program comprising: decoding logic for decoding instructions; and processing logic for selectively applying a vector processing operation specified by the instruction to an input data vector including a plurality of input data items at respective positions within the input data vector, the decoding logic being configured to, in response to a vector combine instruction specifying a plurality of source vector registers, each including a source data element at a plurality of data element positions, one or more further source vector registers, and one or more destination registers, generate control signals that cause the processing logic to, for each data element position of the plurality of data element positions, extract a first source data element from the data element position of each source vector register and extract a second source data element from the one or more further source vector registers and perform a combine operation to generate a result data element, the result data element being computed by combining each element of the first source data element and the second source data element, and storing the result data element in a data element position of the one or more destination registers. [Brief description of the drawings]
[0005] The present technique will be further described, by way of example only, with reference to embodiments thereof illustrated in the accompanying drawings, in which: [Figure 1] 1 illustrates a schematic representation of a processing device according to various configurations of the present technique; [Diagram 2] 2 illustrates generally further details of processing circuitry according to various configurations of the present technique; [Diagram 3] 2 illustrates generally further details of processing circuitry according to various configurations of the present technique; [Figure 4A] 4 illustrates generally further details of a processing unit used to execute a vector combine instruction, according to various configurations of the present technique; [Figure 4B] 4 illustrates generally further details of a processing unit used to execute a vector combine instruction, according to various configurations of the present technique; [Diagram 5]4A-4C illustrate schematic details of vector registers and tile registers according to various configurations of the present technique; [Figure 6] 2 illustrates generally further details of processing circuitry according to various configurations of the present technique; [Figure 7] 2 illustrates generally further details of processing circuitry according to various configurations of the present technique; [Figure 8] 2 illustrates generally further details of processing circuitry according to various configurations of the present technique; [Figure 9] 2 illustrates generally further details of processing circuitry according to various configurations of the present technique; [Figure 10] 2 illustrates generally further details of processing circuitry according to various configurations of the present technique; [Figure 11A] 2 illustrates generally further details of processing circuitry according to various configurations of the present technique; [Figure 11B] 2 illustrates generally further details of processing circuitry according to various configurations of the present technique; [Figure 12] 2 illustrates generally further details of processing circuitry according to various configurations of the present technique; [Figure 13] 1 illustrates a schematic of a sequence of steps performed by a processing device according to various configurations of the present technique; [Figure 14] 1 illustrates a schematic representation of a simulator implementation of a processing device according to various configurations of the present technique;
[0006] Some example configurations provide a processing device comprising a decode circuit for decoding an instruction and a processing circuit for selectively applying a vector processing operation specified by the instruction to an input data vector including a plurality of input data items at respective positions within the input data vector. The decode circuit is configured to generate control signals in response to a vector combine instruction that specifies a plurality of source vector registers, each including a source data element at a plurality of data element positions, one or more further source vector registers, and one or more destination registers. The processing device is configured to generate control signals that cause the processing circuit to, for each data element position of the plurality of data element positions, extract a first source data element from the data element position of each source vector register, extract a second source data element from the one or more further source vector registers, and perform a combine operation to generate a result data element, the result data element being calculated by combining each of the first source data element and the second source data element, and store the result data element in a data element position of the one or more destination registers.
[0007] A processing device may comprise processing circuitry for performing data processing operations on input data vectors. The input data vector has multiple data elements at different positions within the data vector. Some processing devices are configured to perform arithmetic or logical operations to combine elements from the input data vectors and generate result data elements for storage in a result data vector. The inventors have recognized that it is often desirable to perform certain operations repeatedly during certain types of computation. These operations can be provided by executing a number of existing instructions that are performed using common processing circuitry, for example by operating on one source element at a time, but this approach can have a significant impact on performance. Thus, by providing a processing circuitry with specific circuitry tailored to respond to a single vector combine instruction, a particularly efficient processing device can be provided.
[0008] To this end, the arrangement is configured to respond to a vector composition instruction. The vector composition instruction is part of an instruction set architecture that provides a complete set of instructions available to a programmer to interact with the processing circuit. The instructions of the instruction set architecture are decoded by a decoding circuit that functions to interpret the instructions of the instruction set architecture to control the processing circuit to respond to the instruction. Each vector composition instruction specifies a plurality (i.e., two or more) of source vector registers. Each of the plurality of source vector registers is composed of a plurality of data elements stored in a plurality of data element locations. Data elements of different sizes may be provided according to a particular arrangement or may be flexibly adjusted based on the particular arrangement. In some exemplary arrangements, the vector register is a 512-bit vector register that may be configured as, for example, 64 data elements of 8-bit size, 32 data elements of 16-bit size, 16 data elements of 32-bit size, 8 data elements of 64-bit size, 4 data elements of 128-bit size, or 2 data elements of 256-bit size. In other exemplary configurations, the vector register is a 256-bit vector register that may be configured, for example, as 16 data elements of 16-bit size, 32 data elements of 8-bit size, 8 data elements of 32-bit size, or 4 data elements of 64-bit size. These sizes are provided by way of example only, and it will be readily apparent to one of ordinary skill in the art that other vector register sizes may be incorporated into the configurations described herein. The vector combine instruction also specifies one or more further source vector registers, the one or more further source vector registers being specified in addition to the source vector registers. The one or more further source vector registers each comprise a further plurality of data elements at a further plurality of data element positions. The vector combine instruction also specifies one or more destination registers comprised of a plurality of destination data elements at a plurality of destination data element positions. Each of the one or more further source vector registers and each of the one or more destination registers may be configured to have the same number of data elements of the same size as the source registers.Alternatively, one or more further source vector registers and / or one or more destination registers may be configured to hold different numbers of data elements of different sizes for the multiple source registers.
[0009] The decode circuitry is responsive to the vector composition instruction to cause the processing circuitry to perform a sequence of steps for each data element position of the plurality of data element positions. For example, if each source register holds N data elements, the decode circuitry causes the processing circuitry to perform steps for data element positions 0, 1, 2, ..., N-2, N-1. Although the steps are presented sequentially, this is for illustrative purposes only and any of the following steps may be performed in an order different from that specified, or multiple steps may be performed in parallel. The decode circuitry causes the processing circuitry to extract a first set of source data elements. Each of the first source data elements is extracted from a data element position of a respective source vector register. The decode circuitry also causes the processing circuitry to extract a second set of source data elements from one or more further source vector registers. The elements of the first source data elements are extracted such that each data element is extracted from the same data element position of a different one of the plurality of source vector registers. Each data element of the second source data elements is extracted from one or more further source vector registers. However, the location in the further source vector register from which the second source data element is extracted is not so limited and may be flexibly determined based on the implemented configuration. The decode circuit is also configured to control the processing circuit to perform a compositing operation to generate a result data element. The result data element is calculated by combining elements from the first source data element and the second source data element such that the result data element depends on each data element of the first source data element and the second source data element. In this way, each result data element depends on a data element extracted from each of the plurality of source vector registers and on a data element extracted from one or more further source vector registers. The decode circuit also controls the processing circuit to store the result data element in a data processing element location of one or more destination registers. In this way, the number of result data elements generated is equal to the number of elements in each of the plurality of source vector registers.
[0010] The compositing operation is not limited, and various configurations of the compositing operation are described below. In some exemplary configurations, the compositing operation includes a source compositing operation that generates intermediate data elements, each intermediate data element being generated by compositing a corresponding one of the first source data elements with a second source data element, and an intermediate compositing operation that combines the intermediate data elements to generate a result data element. The terms source compositing operation and intermediate compositing operation are used to distinguish the compositing operation performed (e.g., to distinguish the mathematical or logical operation performed). It will be understood by those skilled in the art that these operations may be performed in parallel by the same functional block of a circuit, or may be performed by sequential blocks of a circuit operating sequentially on each other. The source compositing operation combines each corresponding first source data element of the first source data element with a second source data element to generate an intermediate data element. Thus, each intermediate data element depends on one or more elements of the corresponding first source data element and the second source data element. The compositing operation may use each of the second source data elements, or only a subset of the second source data elements. As a result, the number of intermediate data elements is the same as the number of first source data elements. The compositing operation also includes an intermediate compositing operation that generates a result data element by compositing each of the intermediate data elements together. Thus, a single result data element is generated by the intermediate compositing operation (although this is repeated for each data element position of the multiple data element positions in the multiple source vector registers, either sequentially or in parallel).
[0011] In some configurations, the source compositing operation is a multiplication operation, and combining a corresponding one of the first source data elements with a second source data element includes multiplying the corresponding one of the second source data elements with a corresponding second source data element to generate an intermediate data element. The intermediate data element thus includes an element received from the same position of each of the multiple source vectors multiplied by an element of one or more further source vector registers. In these configurations, the elements in the intermediate data element may be represented by the following formula:
[0012]
number
[0013] In some alternative configurations, the source compositing operation is a scaling operation that includes extracting one or more first scaling values from the second source data element, extracting one or more second scaling values from the second source data element, performing one of an addition operation to add a corresponding first scaling value of the one or more first scaling values to a corresponding first source data element of the first source data element to generate a corresponding intermediate scaled element, and a subtraction operation to subtract a corresponding first scaling value from the corresponding first source data element to generate a corresponding intermediate scaled element, and multiplying each of the corresponding intermediate scaled elements by a corresponding second scaling value of the one or more second scalings to generate a corresponding intermediate data element of the intermediate data elements. Thus, the intermediate data elements correspond to a scaling of the first source data element (extracted from the first plurality of source vector registers) based on information stored in the second source data element (and extracted from one or more further source vector registers). In these alternative configurations, the elements in the intermediate data element can be represented by the following formula:
[0014]
number
[0015]
number
[0016] The intermediate compositing operation may be defined in various ways and in some configurations may be combined with a source compositing operation as set forth in equation (1), or in other configurations may be combined with a source compositing operation as set forth in equation (2) or equation (3). In some exemplary configurations, the intermediate compositing operation is an accumulation operation, and combining the intermediate data elements to generate a result data element includes accumulating the intermediate data elements. Thus, for each data element position of the multiple data element positions of the multiple source vector registers, the intermediate compositing operation receives all of the intermediate data elements and accumulates them to generate a single result data element. Mathematically, the accumulation operation may be expressed as follows:
[0017]
number
[0018] In some exemplary configurations, the intermediate data elements are first intermediate data elements, and the intermediate compositing operation includes a first intermediate compositing operation that combines the first intermediate data element to generate a second intermediate data element, and a second intermediate compositing operation that combines the second intermediate data element with a destination data element extracted from a data element location of one or more destination registers. In this manner, the destination registers may be used to store a further set of data elements to be combined with the multiple source vector registers and one or more further source vector registers, thereby increasing the flexibility of the processing unit when responding to vector compositing instructions.
[0019] The first intermediate compositing operation and the second intermediate compositing operation can be defined in various ways. In some example configurations, the first intermediate compositing operation is an accumulation operation, and combining the first intermediate data elements to generate the second intermediate data element includes accumulating the first intermediate elements. In such configurations, the second intermediate data element can be expressed as:
[0020]
number
[0021] The second intermediate compositing operation may be defined in a variety of ways. In some example configurations, the second intermediate compositing operation is one of a masking operation that masks the value of the second intermediate data element, or a multiplication operation or a scaling operation that scales the second intermediate data element. In some configurations, the second intermediate compositing operation is an accumulation operation, and combining the second intermediate data element with the destination data element includes accumulating the second intermediate data element with the destination data element. In such configurations, the result data element may be expressed as:
[0022]
number
[0023] In configurations where the compositing operation is split as a first compositing operation and a second compositing operation, the first processing operation and the second processing operation may be performed sequentially or in parallel. In some configurations, for each data element position, at least a subset of the first compositing operation is performed in parallel with the second compositing operation. The subset of first compositing operations may refer to a subset of each operation such that a portion of each first compositing operation is performed in parallel with a portion of the second compositing operation. Alternatively or additionally, the processing circuitry may be arranged such that the subset of first compositing operations includes a subset of the complete compositing operation that is performed in parallel with the complete second compositing operation. For example, in configurations where the first compositing operation is a multiplication operation and the second compositing operation is an accumulation operation, the first compositing operation and the second compositing operation may be implemented using one or more fused multiply-accumulate circuits to perform the multiplication operation of the first compositing operation in parallel with the compositing operation of the second compositing operation. In this manner, the compositing operations may be implemented in a circuit in a compact and efficient manner.
[0024] In some configurations, the compositing operation includes a dot product operation to generate a dot product of a first source data element and a second source data element as a result data element. In such configurations, the dot product operation may be implemented using any dot product circuit. In some exemplary configurations, the dot product operation may be split into a first compositing operation and a second compositing operation as described above, while in other configurations, the dot product operation may be performed by a single functional circuit that incorporates all circuitry required for the multiplication and addition steps of the dot product operation.
[0025] The result data elements of the one or more destination registers may be variously defined and in some configurations are spread across the one or more destination registers. In some configurations, the size of the result data elements is specified in the vector combine instruction. In some example configurations, the result data element size of each result data element is equal to the source data element size of each source data element. In such configurations, the one or more destination registers is a single destination register that is the same size (number of bits and number of data elements) as each of the multiple source vector registers. In some configurations, the result data element size of each result data element is larger than the source data element size of each source data element. In such configurations, the vector combine instruction is an expansion instruction to widen the number of bits associated with the data elements and the result data elements are spread across the multiple destination registers.
[0026] For example, in some configurations, the source data element size is one of 8 bits, the result data element size is 32 bits, the source data element size is 16 bits, and the result data element size is 64 bits. In some configurations, some of the destination registers of the one or more destination registers are determined based on a ratio of the result data element size to the source data element size. In each of the aforementioned sets of result and source data element sizes, the result data element size is four times as large as the source data element size, and thus the one or more destination registers include four destination registers. In this manner, it is possible to provide a sufficient number of bits in the one or more destination registers to enable a compositing operation to be performed without loss of precision.
[0027] The distribution of the result elements in the destination registers can be defined in various ways. In some exemplary configurations, the one or more destination registers are arranged to form a result array with a number of rows equal to the number of destination registers and a number of columns equal to the number of data elements in each destination register, and the result data elements are arranged in the result array in row-major order. In this way, the result elements can be arranged in the one or more destination registers in the same order as they appear in the source register. In some alternative configurations, the one or more destination registers are arranged to form a result array with a number of rows equal to the number of destination registers and a number of columns equal to the number of data elements in each destination register, and the result data elements are arranged in the result array in column-major order. By arranging the result data elements in the one or more destination registers in this way, the result data elements are stored in the destination register closer to where the source data elements are extracted, thus achieving a more compact design.
[0028] In some exemplary configurations, the processing circuitry uses all of the data elements in the one or more further source vector registers. However, in some configurations, only a subset of the data elements of the one or more further source vector registers are used. The selection of the source elements may be hard-coded into the data processing apparatus. However, in some configurations, the vector compositing instruction specifies a location of the second source data element in the one or more further source vector registers. This may provide improved flexibility and allow the same further source vector register to be used for multiple vector compositing operations. In some configurations, the location specified in the one or more further source vector registers corresponds to a particular location in the one or more further source vector registers. Alternatively, the location refers to a relative location within each of a plurality of subsections of the one or more further source vector registers. This provides a particularly efficient apparatus for performing iterative vector compositing operations in which a different portion of the one or more further source vector registers is used for each operation and the location is specified relative to the location read for a current instance of the operation. In some exemplary configurations, each of the multiple source vector registers, the one or more further source vector registers, and the one or more destination registers may be divided into chunks. For example, each of the registers (including one or more further source vector registers, the plurality of source vector registers, and the destination register) may be divided into four 128-bit chunks, and the location specified in the vector combine instruction identifies one or more data elements to be extracted from within each of the 128-bit chunks as a second source data element to be used for the 128-bit chunk of each of the plurality of source vector registers (e.g., by duplicating one or more of the identified data elements). For example, if the elements are 8-bit and four consecutive data elements are extracted from each 128-bit chunk (out of a total of 16 8-bit elements per chunk), there are four positions within each chunk that may be selected. In this case, the relative location may be set to (for example) a third relative location within each 128-bit chunk.In this case, data elements 8-11 (i.e. from within the third position of the first 128-bit chunk) are selected and applied to a compositing operation associated with a first 128-bit chunk of data elements in the plurality of source vector registers (e.g. by replicating data elements 8-11 extracted from one or more further source vector registers four times or by repeatedly using the same extracted data element), and then a resulting data element from the compositing operation associated with the first 128-bit chunk is stored in a first 128-bit chunk of one or more destination registers. Data elements 24-27 (i.e. from within the third position of the second 128-bit chunk) are selected and applied to a compositing operation associated with a second 128-bit chunk of data elements in the plurality of source vector registers (e.g. by replicating data elements 24-27 extracted from one or more further source vector registers four times or by repeatedly using the same extracted data element), and then a resulting data element from the compositing operation associated with the second 128-bit chunk is stored in a second 128-bit chunk of one or more destination registers. Data elements 40-43 are selected and applied to a compositing operation associated with a third 128-bit chunk of data elements in the plurality of source vector registers (e.g., by replicating data elements 40-43 four times extracted from one or more further source vector registers or by repeatedly using the same data elements), and then a resulting data element from the compositing operation associated with the third 128-bit chunk is stored in a third 128-bit chunk of one or more destination registers. Data elements 56-59 are selected and applied to a compositing operation associated with a fourth 128-bit chunk of data elements in the plurality of source vector registers (e.g., by replicating data elements 56-59 four times extracted from one or more further source vector registers or by repeatedly using the same data elements), and then a resulting data element from the compositing operation associated with the fourth 128-bit chunk is stored in a fourth 128-bit chunk of one or more destination registers.It will be readily apparent to one skilled in the art that a 128-bit size is used as an example, and any chunk size (smaller than, the same as, or larger than the size of one of the one or more additional source vector registers) may be used.
[0029] While the number of source vector registers and the number of source data elements used in the one or more further source vector registers may be variously defined according to any of the aforementioned configurations, in some example configurations the plurality of source vector registers includes two source vector registers and each of the one or more further source vector registers includes two source data elements, while in other example configurations the plurality of source vector registers includes four source vector registers and each of the one or more further source vector registers includes four source data elements.
[0030] The numeric format of each element can be defined in various ways, and in some exemplary configurations, each element of each data vector comprises one of a signed integer value and an unsigned integer value. Furthermore, in some exemplary configurations, each element of the further vector register comprises one of a signed integer value and an unsigned integer value. Thus, different configurations provide any combination of each data vector of the plurality of source vector registers and the further vector register. Thus, in some configurations, each element of each data vector is a signed integer value and each element of the further vector register is a signed integer value, in other configurations, each element of each data vector is a signed integer value and each element of the further vector register is an unsigned integer value, in other configurations, each element of each data vector is an unsigned integer value and each element of the further vector register is a signed integer value, and in other exemplary configurations, each element of each data vector is an unsigned integer value and each data element of the further vector register is an unsigned integer value.
[0031] In some exemplary configurations, the processing circuitry is arranged to generate each result data element for each element position in sequence, resulting in a reduced circuit footprint. In other exemplary configurations, the processing circuitry is configured to generate the result data elements for each data element position in parallel. Generating the result data elements in parallel allows for faster operation of the vector combine instruction and improves scalability.
[0032] As described, the number of second source data elements can be variously defined and specified as part of the vector combine instruction. However, in some configurations, the number of second source data elements extracted from one or more further source vector registers is equal to the number of source registers in the plurality of source registers. This option is particularly useful when performing dot product operations or matrix-vector product calculations.
[0033] In some configurations, the destination register is a vector register. However, in some configurations, the one or more destination registers are one or more horizontal or vertical tile slices of one or more tile registers, each of which contains a vertically and horizontally addressable two-dimensional array of data elements. Conceptually, tile registers are to vector registers, and vector registers are to scalar registers. Tile registers provide two-dimensional arrays of scalar data elements and are particularly efficient for matrix-vector or matrix-matrix computations. Each tile register can be addressed in its entirety or with respect to a vertical or horizontal slice (corresponding to a column or row, respectively) of the tile register. By providing the tile registers as storage destinations, subsequent arithmetic or logical processing operations can be based on the result data elements without requiring further operations to reorder or rearrange the result data elements. Rather, the appropriate row or column (horizontal or vertical tile slice) can be selected from the tile register.
[0034] The manner in which the second source data element is extracted from the one or more further source vector registers may be defined in various ways. In some configurations, the one or more further source vector registers include the same number of vector registers as the plurality of source vector registers, and extracting the second source data element from the one or more further source vector registers includes extracting the second source data element from a data element position of each further source vector register. In such configurations, the one or more further source vector registers are treated in the same manner as the plurality of source vector registers. Thus, for each data element position, the first source data element includes one element for each of the plurality of source vector registers, and the element is extracted from the same position in each of the plurality of source vector registers. Similarly, for each data element position, the second source data element includes one element for each of the one or more source vector registers, and the element extracted from each of the one or more further source vector registers is extracted from the same position in the one or more further source vector registers.
[0035] In some alternative configurations, extracting the second source data elements from the one or more further source vector registers includes extracting the same set of source data elements for each data element position. In such configurations, it may not be necessary to repeatedly perform the step of extracting the second source data elements from the one or more further source vector registers. Rather, the extraction used for the compositing operation at each data element position of the plurality of data element positions may be performed once. In such configurations, the number of the one or more further source vector registers may be defined in a variety of ways. In some configurations, multiple further source vector registers may be defined. In other configurations, the one or more further source vector registers include a single further source vector register. This approach allows for a more compact implementation including fewer vector registers.
[0036] Specific example configurations will now be described with reference to the accompanying drawings.
[0037] Figure 1 shows diagrammatically a processing device 10 in which various examples of the present technique can be embodied. The device comprises a data processing circuit 12 which performs data processing operations on data items in response to a sequence of instructions which it executes. The instructions are fetched from a memory 14 which the data processing device has access to, and for this purpose a fetch circuit 16 is provided in a manner well known to those skilled in the art. Furthermore, the instructions fetched by the fetch circuit 16 are passed to an instruction decode circuit 18 (also called decode circuit) which generates control signals arranged to control various aspects of the configuration and operation of the processing circuit 12 as well as the set of registers 20 and the load / store unit 22. In general, the data processing circuit 12 may be arranged in a pipelined manner, the details of which are not relevant to the present technique. The general arrangement which Figure 1 represents is well known to those skilled in the art and a further detailed description is dispensed with solely for reasons of brevity. As can be seen in Fig. 1, the registers 20 each comprise storage for a number of data elements, such that a processing circuit can apply a data processing operation to a specified data element in the specified register, or to a specified group of data elements ("vectors") in the specified register. In particular, the illustrated data processing apparatus is directed to the performance of vectorized data processing operations, and in particular to the execution of complex processing instructions on data elements held in the registers 20, further description of which follows in more detail below with reference to some specific embodiments. Data values required by the data processing circuit 12 in the execution of instructions, and data values generated as a result of the data processing instructions, are written to and read from the memory 14 by the load / store unit 22. It should also be noted that in general, the memory 14 of Fig. 1 can be viewed as an example of a computer-readable storage medium in which the instructions of the present technique may be stored, typically as part of a sequence of predetermined instructions ("programs") that the processing circuit will subsequently execute. However, the processing circuit may access such programs from a variety of different sources, such as in RAM, in ROM, via a network interface, etc.This disclosure describes various novel instructions that processing circuitry 12 can execute, and the following figures provide further explanation of the nature of these instructions, modifications to data processing circuitry to support execution of these instructions, etc.
[0038] FIG. 2 illustrates in schematic form further details of the processing circuit 30 of FIG. 1 according to some exemplary configurations. In particular, the processing circuit 30 comprises a number of source vector registers, namely source vector register A 32 and source vector register B 34. The processing circuit 30 also comprises a further source vector register 36 (which in some configurations may be one or more further source vector registers). The processing circuit 30 is controlled by control signals generated by the decoding circuit to perform compositing operations 40(A), 40(B), 40(C), 40(D) to generate and store result data elements in one or more destination registers. Each of the compositing operations 40(A), 40(B), 40(C), 40(D) is performed on data elements in one of the data element positions of the source vector register A 32 and the source vector register B 34. Furthermore, each compositing operation 40(A), 40(B), 40(C), 40(D) is based on elements of the further source vector register 36. In particular, a compositing operation 40(A) receives as inputs an element 32(A) from source vector register A 32, an element 34(A) from source vector register B 34, and elements 36(C) and 36(D) from a further source vector register 36. These elements are combined by the compositing operation 40(A) to generate a result data element. A compositing operation 40(B) receives as inputs an element 32(B) from source vector register A 32, an element 34(B) from source vector register B 34, and elements 36(C) and 36(D) from a further source vector register 36. These elements are combined by the compositing operation 40(B) to generate a result data element. A compositing operation 40(C) receives as inputs an element 32(C) from source vector register A 32, an element 34(C) from source vector register B 34, and elements 36(C) and 36(D) from a further source vector register 36. These elements are combined by a combination operation 40(C) to produce a result data element.A compositing operation 40(D) receives as inputs an element 32(D) from source vector register A 32, an element 34(D) from source vector register B 34, and elements 36(C) and 36(D) from a further source vector register 36. These elements are combined by compositing operation 40(D) to produce a result data element.
[0039] 3 shows in schematic form the details of the processing circuitry 48 according to some exemplary configurations in which the compositing operation 40 includes source compositing operations 44, 46 and intermediate compositing operations 42. The source compositing operations 44, 46 each combine a corresponding element of the source vector register A 32 or the source vector register B 34 with an element of the further source register 36. In the illustrated example, the source compositing operations 44, 46 receive as input two elements of the further source register 36. However, this is for illustrative purposes only and those skilled in the art will appreciate that the source compositing operations 44, 46 may each receive any same or different subset of elements from the further source register. The source compositing operations combine data elements extracted from the source vector register A 32, the source vector register B 34 and the further source vector register 36 to generate intermediate data elements. The intermediate data elements are fed to the intermediate compositing operations 42, each of which generates a result data element that is stored in one or more destination registers.
[0040] The processing circuitry 48 performs compositing operations, including source compositing operations 44 and 46 and intermediate compositing operation 42, for each element position in each of source vector register A 32 and source vector register B 34. In particular, if the element position is the lowest position corresponding to data element 32(A) in source vector register A 32 and data element 34(A) in source vector register B 34, then source compositing operation 44(A) combines data element 32(A) of source vector register A 32 with data elements 36(C) and 36(D) from a further source vector register 36. Similarly, source compositing operation 46(A) combines data element 34(A) of source vector register B 34 with data elements 36(D) and 36(C) from a further source vector register 36. The outputs of the source compositing operations 44(A) and 46(A) generate intermediate data elements that are fed to intermediate compositing operation 42(A) to generate a result data element that is stored in one or more destination registers.
[0041] Similarly, when the element position is the second lowest position corresponding to data element 32(B) in source vector register A 32 and data element 34(B) in source vector register B 34, source compositing operation 44(B) combines data element 32(B) of source vector register A 32 with data elements 36(C) and 36(D) from the further source vector register 36. Similarly, source compositing operation 46(B) combines data element 34(B) of source vector register B 34 with data elements 36(D) and 36(C) from the further source vector register 36. The outputs of source compositing operations 44(B) and 46(B) generate intermediate data elements which are fed to intermediate compositing operation 42(B) to generate result data elements which are stored in one or more destination registers.
[0042] Similarly, when the element position is the second most significant position corresponding to data element 32(C) in source vector register A 32 and data element 34(C) in source vector register B 34, source compositing operation 44(C) combines data element 32(C) of source vector register A 32 with data elements 36(C) and 36(D) from the further source vector register 36. Similarly, source compositing operation 46(C) combines data element 34(C) of source vector register B 34 with data elements 36(D) and 36(C) from the further source vector register 36. The outputs of the source compositing operations 44(C) and 46(C) generate intermediate data elements which are fed to intermediate compositing operation 42(C) to generate result data elements which are stored in one or more destination registers.
[0043] Similarly, if the element position is the most significant position corresponding to data element 32(D) in source vector register A 32 and data element 34(D) in source vector register B 34, then source compositing operation 44(D) combines data element 32(D) of source vector register A 32 with data elements 36(C) and 36(D) from a further source vector register 36. Similarly, source compositing operation 46(D) combines data element 34(D) of source vector register B 34 with data elements 36(D) and 36(C) from a further source vector register 36. The outputs of source compositing operations 44(D) and 46(D) generate intermediate data elements which are fed to intermediate compositing operation 42(D) to generate a result data element which is stored in one or more destination registers.
[0044] The compositing operations described above are performed by separate compositing units for each of the least significant, second least significant, second most significant, and most significant positions in source vector register A 32 and source vector register B. However, it will be appreciated by those skilled in the art that a single set of compositing circuit blocks (e.g., source compositing elements 44(A) and 46(A) and a single intermediate compositing operation 42(A)) could be provided, the inputs of which could be fed, for example, through a series of demultiplexers, and the output of intermediate compositing operation 44(A) could be multiplexed into each result element position of one or more destination registers.
[0045] 4A illustrates the use of a processing unit 50 configured to perform a series of dot product operations in response to a vector combine instruction, according to various configurations of the present technique. The processing unit 50 comprises a decode circuit 56 that decodes the instruction and provides control signals to a processing circuit 54. The processing unit 50 further comprises a series of registers 52 that are used to store data vectors. In the illustrated configuration, the data processing unit 50 is used to perform a series of dot product operations that correspond to a matrix-vector multiplication. In particular, the data processing unit 50 is used to calculate the result of multiplying a matrix 58 by a vector 60. Mathematically, this operation is performed by performing a series of dot products to calculate the dot products of each row of the matrix 58 and the vector 60.
[0046] The matrix 58 is stored in a number of source vector registers, including source vector register A 62 and source vector register B 64, such that a first column of the matrix 58 is stored in source vector register A 62 and a second column of the matrix 58 is stored in source vector register B 64. The vector 60 is stored in a single further source vector register 66. In the illustrated embodiment, the two elements of the vector 60 are stored as the two least significant elements of the further source vector register 66. However, this is for illustrative purposes only, and it will be understood by those skilled in the art that any position within the further source vector register may be used interchangeably (and may optionally be specified in the vector combine instruction). The number of source vector registers, including source vector register A 62 and source vector register B 64, as well as the further source vector register, are registers stored as registers 52 in the register storage device of the data processing apparatus 50.
[0047] A data processing apparatus 50 having stored vector registers 52 is responsive to a vector compositing instruction. The vector compositing instruction is received by a decode circuit 56 and causes a processing circuit 54 to perform a series of operations for each data element location of the source vector registers. In this case, the processing circuit performs four series of operations (optionally in parallel), one for each of the four source vector register locations. In the illustrated embodiment, the compositing operation includes a dot product instruction, or a multiplication operation as the source compositing operation, and an accumulation operation as the intermediate compositing operation to generate a result data element. The result data element is twice as wide as the source vector elements (e.g., as defined in the vector compositing instruction), and therefore requires two destination vector registers to provide sufficient storage space for the result data element. In this case, the destination vector registers include a result vector register A 68 containing result data elements 68(A), 68(B) and a result vector register B 70 containing result data elements 70(A) and 70(B).
[0048] The decode circuit 56 controls the processing circuit 54 to decode the first source data element A from the source vector register A 62 and the source vector register B, respectively. 1,1 and A 1,2 and generates a result data element 68(B) by extracting second source data elements b1 and b2 from a further source vector register 66. 1,1 b1+A 1,2 and controls processing circuitry 54 to combine the first source data element with the second source data element by performing a dot product operation to generate a result data element 68(B) in which the value of b2 is stored, which thus comprises the first value of result matrix 72 obtained by multiplying matrix 58 by vector 60.
[0049] Decode circuitry 56 controls processing circuitry 54 to generate result data elements 68(A), 70(B), and 70(A) by performing the same sequence of operations to extract elements from corresponding locations in source vector register A 62 and source vector register B 64. These operations may be performed in parallel or sequentially for each source vector register location. In particular, the first source data element A from source vector register A 62 and source vector register B, respectively, 2,1 and A 2,2 and performing a dot product of the first source element with the previously extracted second source element, 2,1 b1+A 2,2 A result data element 68(A) stored in result vector register A 68 having a value of b2 can be generated by inputting the first source data element A from source vector register A 62 and the first source data element B from source vector register B 63, respectively. 3,1 and A 3,2 and performing a dot product of the first source element with the previously extracted second source element, 3,1 b1+A 3,2A result data element 70(B) stored in result vector register B 70 having a value of b2 can be generated by inputting the first source data element A 62 from source vector register A 62 and the first source data element B 63 from source vector register B 64, respectively. 4,1 and A 4,2 and performing a dot product of the first source element with the previously extracted second source element, 4,1 b1+A 4,2 A result data element 70(A) may be generated that is stored in result vector register A 70 having a value of b2.
[0050] FIG. 4B shows a schematic configuration in which the one or more further source vector registers include two further source vector registers. It will be understood by those skilled in the art that for each of the illustrated configurations, one or more further source vector registers may be specified in the vector combine instruction. In the illustrated configuration, the data processing unit 50 is used to perform a series of combing operations corresponding to combining elements from two matrices. In particular, the data processing unit 50 is used to calculate the result of multiplying the matrix 58 by the matrix 600. Mathematically, this operation is performed by performing a series of dot products to calculate the dot products of each row of the matrix 58 with each row of the matrix 600.
[0051] The matrix 58 is stored in a plurality of source vector registers including source vector register A 62 and source vector register B 64 such that a first column of the matrix 58 is stored in source vector register A 62 and a second column of the matrix 58 is stored in source vector register B 64. The matrix 600 is stored in a plurality of further source vector registers such that a first column of the matrix 600 is stored in further source vector register A 660 and a second column of the matrix 600 is stored in further source vector register B 670. The plurality of source vector registers including source vector register A 62 and source vector register B 64 and the plurality of further source vector registers including further source vector register A 660 and further source vector register B 670 are stored as registers 52 in a register storage device of the data processing apparatus 50.
[0052] A data processing apparatus 50 having stored vector registers 52 is responsive to a vector compositing instruction. The vector compositing instruction is received by a decode circuit 56 and causes a processing circuit 54 to perform a series of operations for each data element location of the source vector registers. In this case, the processing circuit performs four series of operations (optionally in parallel), one for each of the four source vector register locations. In the illustrated embodiment, the compositing operation includes a dot product instruction, or a multiplication operation as the source compositing operation, and an accumulation operation as the intermediate compositing operation to generate a result data element. The result data element is twice as wide as the source vector elements (e.g., as defined in the vector compositing instruction), and therefore requires two destination vector registers to provide sufficient storage space for the result data element. In this case, the destination vector registers include a result vector register A 68 containing result data elements 68(A), 68(B) and a result vector register B 70 containing result data elements 70(A) and 70(B).
[0053] The decode circuit 56 reads the first source data element A from the source vector register A 62 and the source vector register B, respectively. 1,1 and A 1,2and extracts the second source data element B from the further source vector register A 660 and the further source vector register B 670, respectively. 1,1 and B. 1,2 The decryption circuit 56 further controls the processing circuit 54 to generate a result data element 68(B) by extracting 1,1 B 1,1 +A 1,2 B 1,2 The processing circuitry 54 controls the first source data element to combine the second source data element by performing a dot product operation to generate a result data element 68(B) in which a value of x is stored, which thus comprises the first value of a result matrix 72 obtained by multiplying the matrix 58 by the vector 60.
[0054] Decode circuitry 56 controls processing circuitry 54 to generate result data elements 68(A), 70(B), and 70(A) by performing the same sequence of operations to extract elements from corresponding locations in source vector register A 62, source vector register B 64, further source vector register A 660, and further source vector register B 670. These operations may be performed in parallel or sequentially for each source vector register location. In particular, the first source data element A from source vector register A 62 and source vector register B 670, respectively, is generated by: 2,1 and A 2,2 , and extracting the first source element, and the second source element B extracted from the further source vector register A 660 and the further source vector register B 670. 2,1 and B. 2,2 By performing a dot product with A, the processing circuit 2,1 B 2,1 +A 2,2 B 2,2 A result data element 68(A) stored in result vector register A 68 having a value of 3,1 and A 3,2, and extracting the first source element, and the second source element B extracted from the further source vector register A 660 and the further source vector register B 670. 3,1 and B. 3,2 By performing a dot product with A, the processing circuit 3,1 B 3,1 +A 3,2 B 3,2 A result data element 70(B) stored in result vector register B 70 having a value of 4,1 and A 4,2 , and extracting the first source element, and the second source element B extracted from the further source vector register A 660 and the further source vector register B 670. 4,1 and B. 4,2 By performing a dot product with A, the processing circuit 4,1 B 4,1 +A 4,2 B 4,2 may generate a result data element 70(A) stored in result vector register B 70 having a value of:
[0055] 5 illustrates in schematic form further details of registers provided for a data processing device 80 according to various exemplary configurations. The processing device 80 comprises a vector register store 82 and a tile register store 86, in addition to the decoding circuitry 92 and processing circuitry 90 described above. The vector register store 82 comprises N vector registers 84(1), 84(2), ..., 84(N). The tile register store 86 comprises tile registers 88(1), ..., 88(M). The number of vector and tile registers provided may vary depending on the configuration. Each vector register contains multiple elements and may be addressed on a vector register basis or a vector element basis. The tile registers 88 are arranged as a two-dimensional array of elements. Each tile register may be addressed on a tile basis, an element basis, or with respect to rows (horizontal slices) or columns (vertical slices) of tile registers.
[0056] 6 illustrates the operation of a processing circuit according to various exemplary configurations of the present technique. In the illustrated configuration, the compositing operation includes a first compositing operation which is a multiplication operation 96, 98, and a second compositing operation which is an addition operation 100. The processing circuit is configured to combine a number of source vector registers including a source vector register A 90 and a source vector register B 92 with a further source vector register 94 in response to a vector compositing instruction. The processing circuit is configured to output a result element to a result vector register A 102 which contains the same number of elements of the same size as each of the source vector registers 90, 92. In operation, the processing circuit performs a multiplication operation 98 to combine an element of the source vector register A 90 with an element b1 of the further source vector register 94, and a multiplication operation 96 to combine an element of the source vector register B 92 with an element b2 of the further source vector register 94. The results of the corresponding sets of multiplication operations are combined via an accumulation operation 100 to generate a result data element which is stored in the result vector register A 102.
[0057] The processing circuit of Figure 6 provides circuitry for performing four sets of combing operations (identified by a letter included in brackets at the end of the reference number), each associated with one of the result data elements. The most significant (left-most) result data element is generated from the most significant elements of source vector register A 90 and source vector register B 92 combined with corresponding elements of further source vector register 94 using multiplication operations 96(D) and 98(D) and an accumulation operation 100(D) for combining the outputs of multiplication operations 96(D) and 98(D). Similarly, the second most significant elements of result vector register 102 are generated from the second most significant elements of source vector register A 90 and source vector register B 92 combined with corresponding elements of further source vector register 94 using multiplication operations 96(C) and 98(C) and an accumulation operation 100(C) for combining the outputs of multiplication operations 96(C) and 98(C). Similarly, the second lowest element of result vector register 102 is generated from the second lowest elements of source vector register A 90 and source vector register B 92 combined with corresponding elements of further source vector register 94 using multiplication operations 96(B) and 98(B) and an accumulation operation 100(B) to combine the outputs of multiplication operations 96(B) and 98(B). Finally, the lowest (rightmost) result data element is generated from the lowest elements of source vector register A 90 and source vector register B 92 combined with corresponding elements of further source vector register 94 using multiplication operations 96(A) and 98(A) and an accumulation operation 100(A) to combine the outputs of multiplication operations 96(A) and 98(A).
[0058] FIG. 7 shows in schematic form the details of the operations performed by the processing circuitry according to various exemplary configurations. The processing circuitry of FIG. 7 includes the same source compositing operation as described in relation to FIG. 6. The intermediate compositing operation of FIG. 6 includes a first intermediate compositing operation 100 that is identical to the intermediate compositing operation 100 of FIG. 6. In addition, the processing circuitry of FIG. 7 is provided with a second intermediate compositing operation 104, which in the illustrated embodiment is an accumulation operation. The second intermediate compositing operation combines the output from the first intermediate compositing operation 100 with the values of data elements already present in the destination vector register 102. The result data elements of the result vector register 102 are the same size as the data elements of the source vector registers 90, 92, respectively.
[0059] 7 provides circuitry for performing four sets of combing operations each associated with one of the result data elements. The lowest result data element is generated from the lowest elements of source vector register A 90 and source vector register B 92 combined with corresponding elements of a further source vector register 94 using multiplication operations 96(D) and 98(D), an accumulation operation 100(D) for combining the outputs of multiplication operations 96(D) and 98(D), and an accumulation operation 104(D) for accumulating the output of accumulation operation 100(D) and an existing value in result vector register 102. Similarly, the second lowest element of result vector register 102 is generated from the second lowest elements of source vector register A 90 and source vector register B 92 combined with the corresponding element of a further source vector register 94 using multiplication operations 96(C) and 98(C), an accumulation operation 100(C) for combining the outputs of multiplication operations 96(C) and 98(C), and an accumulation operation 104(C) for accumulating the output of accumulation operation 100(C) and the existing value in result vector register 102. Similarly, the second most significant element of the result vector register 102 is generated from the second most significant elements of the source vector register A 90 and the source vector register B 92 combined with the corresponding elements of the further source vector register 94 using multiplication operations 96(B) and 98(B), an accumulation operation 100(B) for combining the outputs of the multiplication operations 96(B) and 98(B), and an accumulation operation 104(B) for accumulating the output of the accumulation operation 100(B) and the existing value in the result vector register 102. Finally, the most significant result data element is generated from the most significant elements of the source vector register A 90 and the source vector register B 92 combined with the corresponding elements of the further source vector register 94 using multiplication operations 96(A) and 98(A), an accumulation operation 100(A) for combining the outputs of the multiplication operations 96(A) and 98(A), and an accumulation operation 104(A) for accumulating the output of the accumulation operation 100(A) and the existing value in the result vector register 102.The accumulation operations 100 and 104 and the multiplication operations 96 and 98 may be provided as separate separate logic blocks or a single combined circuit that performs the accumulation and multiplication steps as described with reference to the accumulation operations 100 and 104 and the multiplication operations 96 and 98, respectively. Furthermore, a single complete set of accumulation and multiplication circuits, e.g., multiplication circuits 96(A) and 98(A) and accumulation units 100(A) and 104(A), may be provided to sequentially perform the above-mentioned steps with operands from source vector register A and source vector register B selected using switches (demultiplexers) and result data elements provided to corresponding positions in the destination register using switches (multiplexers).
[0060] FIG. 8 illustrates in schematic detail the operations performed by the processing circuitry according to various exemplary configurations. The processing circuitry of FIG. 8 includes a source compositing operation and an intermediate compositing operation as described in connection with FIG. 6. The intermediate compositing operation may similarly include a second intermediate compositing operation as described in connection with FIG. 7. The processing circuitry is configured to generate result data elements that are twice as wide as the data elements of the input source registers 90, 92. The result data elements are stored in destination register A 106 and destination register B 108. The destination registers A 106 and destination register B 108 are arranged to form a result array that includes a number of rows equal to the number of destination registers and a number of columns equal to the number of data elements in each destination register. Furthermore, the processing circuitry is configured to arrange the result data elements in the result array in row-major order. In particular, element A of source register B 92 is 1,2 and A 2,2 and element A of source register A90 1,1 and A 2,1 The result data element associated with is stored in destination register A 106. Similarly, element A of source register B 92 3,2 and A 4,2 and element A of source register A90 3,1 and A 4,1 The result data element associated with is stored in destination register B108.
[0061] FIG. 9 illustrates in schematic detail the operations performed by the processing circuitry according to various exemplary configurations. The processing circuitry of FIG. 9 includes the same source and intermediate compositing operations as described in connection with FIGS. 6 and 8. The processing circuitry is configured to generate result data elements that are twice as wide as the data elements of the input source registers 90, 92. The result data elements are stored in destination register A 106 and destination register B 108. The destination registers A 106 and destination register B 108 are arranged to form a result array that includes a number of rows equal to the number of destination registers and a number of columns equal to the number of data elements in each destination register. Furthermore, the processing circuitry is configured to arrange the result data elements in the result array in column-major order. In particular, element A of source register B 92 is 1,2 and A 3,2 and element A of source register A90 1,1 and A- 3,1 The result data element associated with is stored in destination register A 106. Similarly, element A of source register B 92 2,2 and A 4,2 and element A of source register A90 2,1 and A 4,1 The result data element associated with is stored in destination register B108.
[0062] FIG. 10 illustrates in schematic form the details of the operations performed by the processing circuitry according to various exemplary configurations. The processing circuitry of FIG. 10 includes source and intermediate compositing operations that are functionally equivalent to the source and intermediate compositing operations described in relation to FIG. 6, FIG. 8 and FIG. 9. It should be noted that in some configurations, the intermediate compositing operations of FIG. 10 may also be modified to incorporate the second intermediate compositing operation described in relation to FIG. 7. The processing circuitry of FIG. 10 differs from the previous embodiment in that at least a subset of the source compositing operations are performed in parallel with the intermediate compositing operations. In particular, the source compositing operations are multiplication operations and the intermediate compositing operations are accumulation operations. The source compositing operation 96 associated with the source vector register B 92 is performed as a multiplication operation that receives two inputs, one input from the source vector register B 92 and one input from the further source vector register 94. The source compositing operation and the intermediate compositing operation associated with the source vector register A 90 are each performed in parallel by a fused multiply-accumulate (FMA) unit 110. Each fused multiply accumulate unit 110 comprises circuitry for performing multiplication 112 and circuitry for performing accumulation 114. The fused multiply accumulate unit 110(A) receives inputs from the source compositing operation 96(A) and the source vector A 90 and outputs a result data element to the destination register A 106. The fused multiply accumulate unit 110(B) receives inputs from the source compositing operation 96(B) and the source vector A 90 and outputs a result data element to the destination register A 106. The fused multiply accumulate unit 110(C) receives inputs from the source compositing operation 96(C) and the source vector A 90 and outputs a result data element to the destination register B 108. The fused multiply accumulate unit 110(D) receives inputs from the source compositing operation 96(D) and the source vector A 90 and outputs a result data element to the destination register B 108. In the illustrated embodiment, the result elements are output in row-major order as described in connection with FIG. 8.However, those skilled in the art will appreciate that the data elements may be output in column-major order to destination registers 106, 108, which may be configured to store result data elements of any size, and that the intermediate compositing operation may also include a second intermediate compositing operation that accumulates the result in the destination register 106, 108.
[0063] 11A and 11B show schematic details of a processing unit according to the present technique in which a scaling operation is performed, with Fig. 11A showing the case where a single set of scaling values is specified in the instruction, and Fig. 11B showing the case where one set of scaling values is provided for each input vector register.
[0064] 11A shows in schematic form the details of the operations performed by the processing circuitry according to various configurations of the present technique. In the illustrated embodiment, the processing circuitry includes a source compositing operation and an intermediate compositing operation, the source compositing operation being a scaling operation including an add / subtract operation 116, 118 and a multiplication operation 120, 122. An operation is performed for each data element position in a number of source vector registers, namely source vector register A 90 and source vector register B 92. Each set of operations (identified by the same letter in brackets at the end of the reference number) includes extracting a first source data element from each position of the number of source vector registers and extracting a second source data element from the further source vector register 94. The operation includes extracting a first scaling value (b2 in this case) from the second source data element extracted from the further source vector register 94. The first scaling value is input to the add / subtract units 116, 118 together with one of the first source data elements. Each first source data element is scaled through addition / subtraction of a first scaling value by the add / subtract units 116, 118. The operation further includes extracting a second scaling value from a second source data element (in this case b1). The results of the add / subtract operations 116, 118 are provided to multiplication units 120, 122, which multiply the results of the add / subtract operations by the second scaling value to generate intermediate data elements. The intermediate data elements are provided to a combination unit 124, which performs an intermediate combination operation to combine the outputs of the multiplication units 120, 122. The intermediate combination operation performed may be any of the operations described herein. For example, the intermediate combination operation may be a multiplication, an accumulation operation, or another mathematical or logical operation. The output from the accumulation operation 124 is stored in the destination register 106, 108.
[0065] Element A from the lowest position of source vector register A90 4,2is combined with a first scaling value b2 using an add / subtract unit 116 (D) to generate an intermediate scaled element. The intermediate scaled element is combined with a second scaling value b1 using a multiplication unit 120 (D) to generate an intermediate data element. Similarly, element A from the lowest position of source vector register B 92 is 4,1 is combined with the first scaling value b2 using addition / subtraction unit 118(D) to generate an intermediate scaled element. The intermediate scaled element is combined with the second scaling value b1 using multiplication unit 122(D) to generate an intermediate data element. The intermediate data elements output by multiplication unit 120(D) and multiplication unit 122(D) are combined through a combination operation 124(D) to generate a result data element stored in destination register B 108.
[0066]
number
[0067]
number
[0068] Element A from the lowest position of source vector register A90 3,2 is combined with a first scaling value b2 using an add / subtract unit 116 (C) to generate an intermediate scaled element. The intermediate scaled element is combined with a second scaling value b1 using a multiplication unit 120 (C) to generate an intermediate data element. Similarly, element A from the lowest position of source vector register B 92 is 3,1is combined with a first scaling value b2 using an addition / subtraction unit 118(C) to generate an intermediate scaled element. The intermediate scaled element is combined with a second scaling value b1 using a multiplication unit 122(C) to generate an intermediate data element. The intermediate data elements output by multiplication unit 120(C) and multiplication unit 122(C) are combined through a combination operation 124(C) to generate a result data element that is stored in destination register B 108.
[0069]
number
[0070] Element A from the lowest position of source vector register A90 2,2 is combined with a first scaling value b2 using an add / subtract unit 116(B) to generate an intermediate scaled element. The intermediate scaled element is combined with a second scaling value b1 using a multiplication unit 120(B) to generate an intermediate data element. Similarly, element A from the lowest position of source vector register B 92 is 2,1 is combined with a first scaling value b2 using addition / subtraction unit 118(B) to generate an intermediate scaled element. The intermediate scaled element is combined with a second scaling value b1 using multiplication unit 122(B) to generate an intermediate data element. The intermediate data elements output by multiplication unit 120(B) and multiplication unit 122(B) are combined through a combination operation 124(B) to generate a result data element that is stored in destination register A 108.
[0071]
number
[0072] Element A from the lowest position of source vector register A90 1,2is combined with a first scaling value b2 using an add / subtract unit 116(A) to generate an intermediate scaled element. The intermediate scaled element is combined with a second scaling value b1 using a multiplication unit 120(A) to generate an intermediate data element. Similarly, element A1,1 from the lowest position of source vector register B 92 is combined with a first scaling value b2 using an add / subtract unit 118(A) to generate an intermediate scaled element. The intermediate scaled element is combined with a second scaling value b1 using a multiplication unit 122(A) to generate an intermediate data element. The intermediate data elements output by multiplication unit 120(A) and multiplication unit 122(A) are combined through a combination operation 124(A) to generate a result data element that is stored in destination register A 108.
[0073]
number
[0074] In the configuration shown in Figure 11A, each of source vector register A 90 and source vector register B is scaled by the same first scaling value and the same second scaling value. In an alternative configuration shown in Figure 11B, source vector register A 90 is scaled by a corresponding first scaling value b4 and a corresponding second scaling value b3, and source vector register B 92 is scaled by a corresponding first scaling value b2 and a corresponding second scaling value b1. 1, b2, b3, and b4 are each stored in a further source vector register 94. In an alternative configuration, the first scaling values b2 and b4 can be stored in one further source vector register and the second scaling values b1 and b3 can be stored in a different further source vector register. In another configuration, each of the four scaling values can be stored in a separate source vector register. The operation of the circuit shown in Figure 11B is similar to the circuit described in connection with Figure 11A.
[0075] In FIG. 11B, element A from the lowest position of source vector register A90 4,2 is combined with a corresponding first scaled value b2 using an add / subtract unit 116 (D) to generate an intermediate scaled element. The intermediate scaled element is combined with a corresponding second scaled value b1 using a multiplier unit 120 (D) to generate an intermediate data element. Similarly, element A from the lowest position of source vector register B 92 4,1 is combined with a corresponding first scaled value b4 using addition / subtraction unit 118(D) to generate an intermediate scaled element. The intermediate scaled element is combined with a corresponding second scaled value b3 using multiplication unit 122(D) to generate an intermediate data element. The intermediate data elements output by multiplication unit 120(D) and multiplication unit 122(D) are combined through a combination operation 124(D) to generate a result data element that is stored in destination register B108.
[0076]
number
[0077]
number
[0078] Element A from the lowest position of source vector register A90 3,2 is combined with a corresponding first scaling value b2 using an add / subtract unit 116 (C) to generate an intermediate scaled element. The intermediate scaled element is combined with a corresponding second scaling value b1 using a multiplication unit 120 (C) to generate an intermediate data element. Similarly, element A from the lowest position of source vector register B 92 3,1is combined with a corresponding first scaled value b4 using addition / subtraction unit 118(C) to generate an intermediate scaled element. The intermediate scaled element is combined with a corresponding second scaled value b3 using multiplication unit 122(C) to generate an intermediate data element. The intermediate data elements output by multiplication unit 120(C) and multiplication unit 122(C) are combined through a combination operation 124(C) to generate a result data element that is stored in destination register B 108.
[0079]
number
[0080] Element A from the lowest position of source vector register A90 2,2 is combined with a corresponding first scaling value b2 using an add / subtract unit 116(B) to generate an intermediate scaled element. The intermediate scaled element is combined with a corresponding second scaling value b1 using a multiplication unit 120(B) to generate an intermediate data element. Similarly, element A from the lowest position of source vector register B 92 2,1 is combined with a corresponding first scaled value b4 using addition / subtraction unit 118(B) to generate an intermediate scaled element. The intermediate scaled element is combined with a corresponding second scaled value b3 using multiplication unit 122(B) to generate an intermediate data element. The intermediate data elements output by multiplication unit 120(B) and multiplication unit 122(B) are combined through a combination operation 124(B) to generate a result data element that is stored in destination register A 108.
[0081]
number
[0082] Element A from the lowest position of source vector register A90 1,2is combined with a corresponding first scaled value b2 using an add / subtract unit 116(A) to generate an intermediate scaled element. The intermediate scaled element is combined with a corresponding second scaled value b1 using a multiplier unit 120(A) to generate an intermediate data element. Similarly, element A1,1 from the lowest position of source vector register B 92 is combined with a corresponding first scaled value b2 using an add / subtract unit 118(A) to generate an intermediate scaled element. The intermediate scaled element is combined with a corresponding second scaled value b1 using a multiplier unit 122(A) to generate an intermediate data element. The intermediate data elements output by multiplier unit 120(A) and multiplier unit 122(A) are combined through a combine operation 124(A) to generate a result data element that is stored in destination register A 108.
[0083]
number
[0084] Fig. 12 illustrates in schematic form further details of the operations performed by the processing circuitry according to various configurations of the present technique. The vector compositing operation of Fig. 12 combines data elements from multiple source vector registers 130 and a further source vector register 132. Each source vector register 130 is combined with a different element of the further source vector register 132 using a compositing circuit 136. In this case the compositing operation is a multiplication operation, resulting in an intermediate data element 134. Source vector register A 130(A) is multiplied by the most significant (left-most) element of the further source vector register 132 using a compositing circuit 136(A) to generate intermediate data element 134(A). Source vector register B 130(B) is multiplied by the second most significant element of the further source vector register 132 using a compositing circuit 136(B) to generate intermediate data element 134(B). Source vector register C 130 (C) is multiplied by the second least significant element of a further source vector register 132 using a combiner circuit 136 (C) to generate intermediate data element 134 (C). Source vector register D 130 (D) is multiplied by the least significant (rightmost) element of a further source vector register 132 using a combiner circuit 136 (D) to generate intermediate data element 134 (D). The intermediate data elements may be elements of the same width as the elements of the multiple source vector registers or may be wider than the source data elements.
[0085] The intermediate data elements 134 are combined using an intermediate combining circuit 140. In the illustrated embodiment, the combination operation is an accumulation operation. The intermediate combining circuit 140(A) combines the most significant element of each set of intermediate data elements 134 to generate a result data element 142(A). The intermediate combining circuit 140(B) combines the second most significant element of each set of intermediate data elements 134 to generate a result data element 142(B). The intermediate combining circuit 140(C) combines the second least significant element of each set of intermediate data elements 134 to generate a result data element 142(C). The intermediate combining circuit 140(D) combines the least significant element of each set of intermediate data elements to generate a result data element 142(D). The result data elements thus generated correspond to a series of dot product operations performed between vectors formed from elements at the same element position in the source vector register 130 and the further source vector register 132. The result data elements are stored in a result array 142, which may be an array of vector registers or an array of horizontal or vertical slices of one or more tile registers.
[0086] FIG. 13 shows a schematic of a sequence of steps performed by a processing device in response to a sequence of instructions. The flow starts at step S100, where a next instruction is received by the decode circuitry. The flow then proceeds to step S102, where it is determined whether the received instruction is a vector synthesis instruction specifying a plurality of source vector registers, one or more further source vector registers, and one or more destination registers. If it is determined that the instruction is not a vector synthesis instruction, an instruction is issued by the decode circuitry based on the requirements of the particular instruction and the instruction set architecture to which it belongs, before the flow returns to step S100 to await the next instruction. If in step S102 it is determined that the instruction is a vector synthesis instruction, the flow proceeds to step S104, where for each position i of the source vector, a sequence of steps is performed (optionally in parallel). The flow proceeds (for each position i) to step S106, where a first source data element is extracted from each position i of the plurality of source vector registers. The flow then proceeds to step S108, where a second source data element is extracted from the one or more further source vector registers. Flow then proceeds to step S110 where a compositing operation is performed to generate a result data element. The compositing operation involves compositing each data element of the first source data element and the second source data element. Flow then proceeds to step S112 where the result data element is stored in location i of one or more destination vector registers. The method then returns to step S100 to await the next instruction.
[0087] FIG. 14 illustrates a simulator implementation that may be used. While the above embodiments implement the invention in terms of apparatus and methods for operating specific processing hardware supporting the technique, it is also possible to provide an instruction execution environment according to the embodiments described herein implemented by the use of a computer program. Such a computer program is often referred to as a simulator insofar as the computer program provides a software-based implementation of a hardware architecture. Various simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on a host hardware 515, which optionally runs a host operating system (host OS) 510, which supports a simulator program 505. In some arrangements, there may be multiple layers of simulation between the hardware and the instruction execution environment provided, and / or there may be multiple different instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide a simulator implementation that runs at a reasonable speed, but such an approach may be justified in certain situations, such as when it is desired to run code native to another processor for compatibility or reuse reasons. For example, a simulator implementation may provide an instruction execution environment that has additional functionality not supported by the host processor hardware, or that is typically associated with a different hardware architecture. An overview of simulation is given in "Some Efficient Architecture Simulation Techniques", Robert Bedichek, Winter 1990 USENIX Conference, pp. 53-63.
[0088] While embodiments have been described above with reference to particular hardware configurations or functions, equivalent functionality in that respect may be provided in a simulated embodiment by appropriate software configurations or functions. For example, particular circuit configurations may be implemented as computer program logic in a simulated embodiment. Similarly, memory hardware such as registers or caches may be implemented as software data structures in a simulated embodiment. In arrangements in which one or more of the hardware elements referenced in the foregoing embodiments reside in host hardware (e.g., host processor 730), some simulated embodiments may use the host hardware where suitable.
[0089] The simulator program 505 may be stored on a computer-readable storage medium (which may be a non-transitory medium) and provides a program interface (an instruction execution environment) to the target code 500 that is the same as an application program interface of the hardware architecture being modeled by the simulator program 505. The simulator program 505 comprises decode logic 520 that decodes instructions and processing logic 530 that selectively applies vector processing operations specified by the instructions to an input data vector that includes multiple input data items at respective positions within the input data vector. The decode logic 520 is configured to generate control signals in response to a vector composition instruction specifying a plurality of source vector registers each containing a source data element in a plurality of data element locations, one or more further source vector registers, and one or more destination registers to cause the processing logic 530 to perform a composition operation to extract, for each data element location of the plurality of data element locations, a first source data element from the data element location of each source vector register, extract a second source data element from the one or more further source vector registers, and generate a result data element, the result data element being calculated by combining each element of the first source data element and the second source data element, and store the result data element in a data element location of the one or more destination registers. Thus, the program instructions of the target code 500, including the complex processing instructions discussed above, may be executed from within an instruction execution environment using the simulator program 505, such that a host hardware 515 that does not actually have the hardware characteristics of the above-mentioned devices can emulate these characteristics.
[0090] In summary, a processing apparatus, method and computer program are provided. The apparatus comprises a decode circuit for decoding an instruction and a processing circuit for applying a vector processing operation specified by the instruction. The decode circuit is configured to, in response to a vector combine instruction specifying a plurality of source vector registers each including source data elements at a plurality of data element locations, one or more further source vector registers and one or more destination registers, cause the processing circuit to, for each data element location, extract a first source data element from the data element location of each source vector register, extract a second source data element from the one or more further source vector registers, generate a result data element by combining each of the first source data element and the second source data element, and store the result data element in a data element location of the one or more destination registers.
[0091] In this application, the term "configured to..." is used to mean that an element of an apparatus has a configuration capable of performing a defined operation. In this context, "configuration" refers to a method of arrangement or interconnection of hardware or software. For example, an apparatus may have dedicated hardware that provides the defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not imply that an apparatus element needs to be modified in any way to provide the defined operation.
[0092] Although illustrative embodiments have been described in detail herein with reference to the accompanying drawings, it will be understood that the invention is not limited to exact embodiments thereof, and various changes, additions and modifications may be made by those skilled in the art without departing from the scope and spirit of the invention as defined in the appended claims. For example, various combinations of the features of the following dependent claims may be made with the features of the independent claims without departing from the scope of the invention.
Claims
1. A processing device, comprising a decoding circuit for decoding instructions, and a processing circuit for selectively applying the vector processing operation specified by the instruction to an input data vector including a plurality of input data items at respective positions within the input data vector, wherein, in response to a vector composition instruction in which the decoding circuit designates a plurality of source vector registers each including a source data element at a plurality of data element positions, one or more further source vector registers, and one or more destination registers, the processing circuit, for each data element position of the plurality of data element positions, extracts a first source data element from the data element position of each source vector register, extracts a second source data element from the one or more further source vector registers, executes a composition operation to generate a result data element, the result data element being calculated by composing each element of the first source data element and the second source data element, and is configured to generate a control signal for storing the result data element in the data element position of the one or more destination registers. A processing device.
2. The composition operation is a source composition operation for generating an intermediate data element, each intermediate data element being generated by composing a corresponding first source data element among the first source data elements with the second source data element, and an intermediate composition operation for composing the intermediate data elements to generate the result data element. The processing device according to claim 1.
3. The source composition operation is a multiplication operation, Synthesizing the corresponding first source data element among the first source data elements with the second source data element includes multiplying the corresponding first source data element among the corresponding first source data elements and the corresponding second source data element among the second source data elements to generate the intermediate data element. The processing apparatus according to claim 2.
4. The source synthesis operation extracting one or more first scaling values from the second source data element; extracting one or more second scaling values from the second source data element; performing one of an addition operation of adding the corresponding first scaling value among the one or more first scaling values to the corresponding first source data element among the first source data elements to generate the corresponding intermediate scaled element, and a subtraction operation of subtracting the corresponding first scaling value from the corresponding first source data element to generate the corresponding intermediate scaled element; The processing apparatus according to claim 2, which is a scaling operation including multiplying each of the corresponding intermediate scaled elements by the corresponding second scaling value among the one or more second scaling values to generate the corresponding intermediate data element among the intermediate data elements.
5. The intermediate synthesis operation is a cumulative operation, Synthesizing the intermediate data elements to generate the result data element includes accumulating the intermediate data elements. The processing apparatus according to claim 2.
6. The intermediate data element is a first intermediate data element, and the intermediate synthesis operation a first intermediate synthesis operation of synthesizing the first intermediate data element to generate a second intermediate data element; The processing apparatus according to claim 2, further comprising a second intermediate synthesis operation for synthesizing the second intermediate data element with a destination data element extracted from the data element position of the one or more destination registers.
7. The first intermediate synthesis operation is an accumulation operation, Synthesizing the first intermediate data element to generate the second intermediate data element includes accumulating the first intermediate element. The processing apparatus according to claim 6.
8. The second intermediate synthesis operation is an accumulation operation, Synthesizing the second intermediate data element with the destination data element includes accumulating the second intermediate data element with the destination data element. The processing apparatus according to claim 6.
9. For each data element position, at least a subset of the source synthesis operations is executed in parallel with the intermediate synthesis operation, the processing apparatus according to claim 2.
10. The synthesis operation includes a dot product operation for generating a dot product of the first source data element and the second source data element as the result data element, the processing apparatus according to claim 1.
11. The result data element size of each result data element is equal to the source data element size of each source data element, the processing apparatus according to any one of claims 1 to 10.
12. The result data element size of each result data element is larger than the source data element size of each source data element, the processing apparatus according to any one of claims 1 to 10.
13. The source data element size is 8 bits and the result data element size is 32 bits, or The source data element size is 16 bits and the result data element size is 64 bits, The processing device according to claim 12, being any one of them.
14. The processing device according to claim 11, wherein the number of destination registers among the one or more destination registers is determined based on a ratio between the result data element size and the source data element size.
15. The one or more destination registers are arranged to form a result array including a number of rows equal to the number of destination registers and a number of columns equal to the number of data elements in each destination register, and result data elements are arranged in the result array in row-major order. The processing device according to claim 14.
16. The one or more destination registers are arranged to form a result array including a number of rows equal to the number of destination registers and a number of columns equal to the number of data elements in each destination register, and result data elements are arranged in the result array in column-major order. The processing device according to claim 14.
17. The processing device according to any one of claims 1 to 10, wherein the vector synthesis instruction specifies a location of the second source data element in the one or more additional source vector registers.
18. The processing device according to any one of claims 1 to 10, wherein the plurality of source vector registers includes two source vector registers, and each of the one or more additional source vector registers includes two source data elements.
19. The processing device according to any one of claims 1 to 10, wherein the plurality of source vector registers includes four source vector registers, and each of the one or more additional source vector registers includes four source data elements.
20. Each element of each data vector is a signed integer value, and The processing device according to any one of claims 1 to 10, including one of a signed integer value and an unsigned integer value.
21. Each element of the one or more additional vector registers includes a signed integer value and an unsigned integer value, and the processing device according to any one of claims 1 to 10.
22. The processing device according to any one of claims 1 to 10, wherein the processing circuit is configured to generate the result data elements in parallel for each data element position.
23. The number of second source data elements extracted from the one or more additional source vector registers is equal to the number of source registers in the plurality of source registers, and the processing device according to any one of claims 1 to 10.
24. The one or more destination registers are one or more horizontal or vertical tile slices of one or more tile registers, and each of the one or more tile registers includes a two-dimensional array addressable vertically and horizontally by data elements. The processing device according to any one of claims 1 to 10.
25. The one or more additional source vector registers include the same number of vector registers as the plurality of source vector registers, extracting the second source data element from the one or more additional source vector registers includes extracting the second source data element from the data element positions of each additional source vector register, The processing device according to any one of claims 1 to 10.
26. Extracting the second source data element from the one or more additional source vector registers includes extracting the same set of source data elements for each data element position, and the processing device according to any one of claims 1 to 10.
27. The processing apparatus according to claim 26, wherein the one or more additional source vector registers include a single additional source vector register. **Claim 28** A method for operating a processing apparatus comprising a decoding circuit that decodes instructions and a processing circuit that selectively applies a vector processing operation specified by the instructions to an input data vector including a plurality of input data items at respective positions within the input data vector, the method comprising: in response to a vector composition instruction that uses the decoding circuit to specify a plurality of source vector registers each including a source data element at a plurality of data element positions, one or more additional source vector registers, and one or more destination registers, causing the processing circuit to, for each data element position of the plurality of data element positions, extracting a first source data element from the data element position of each source vector register; extracting a second source data element from the one or more additional source vector registers; performing a composition operation to generate a result data element, the result data element being calculated by composing each element of the first source data element and the second source data element; generating control signals for causing the result data element to be stored at the data element positions of the one or more destination registers. **Claim 29** A computer program for controlling a host processing apparatus to provide an instruction execution environment, the computer program comprising: decoding logic for decoding instructions; processing logic for selectively applying a vector processing operation specified by the instructions to an input data vector including a plurality of input data items at respective positions within the input data vector; In response to a vector composition instruction that specifies a plurality of source vector registers each containing source data elements at a plurality of data element positions, one or more additional source vector registers, and one or more destination registers, the decoding logic causes the processing logic to, for each data element position of the plurality of data element positions, extract a first source data element from the data element position of each source vector register, extract a second source data element from the one or more additional source vector registers, execute a composition operation to generate a result data element, the result data element being calculated by composing each element of the first source data element and the second source data element, and generate a control signal to cause the result data element to be stored at the data element position of the one or more destination registers. A computer program configured to