Techniques for performing outer product operations
By employing a processing circuit that performs multiple outer product operations using sub-vectors within source vector operands, the inefficiencies in existing data processing systems are addressed, resulting in improved efficiency and throughput.
Patent Information
- Application Number
- JP2024571940
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-13
- Filing Date
- 2023-05-23
- Publication Date
- 2025-06-26
AI Technical Summary
Existing data processing systems face inefficiencies in performing outer product operations due to suboptimal utilization of array storage devices and processing circuitry resources.
The implementation of a processing circuit with an instruction decoder that performs multiple outer product operations using sub-vectors within source vector operands, storing results in a two-dimensional array within an array storage device.
This approach enhances processing efficiency and throughput by enabling multiple outer product operations to be performed in response to a single instruction, while optimizing the use of memory elements and hardware resources.
Smart Images

Figure 2025519452000001_ABST
Abstract
Description
Technical Field
[0001] This technology relates to the field of data processing, and more specifically, to the processing ability of outer product operations.
[0002] Some modern data processing systems may provide an array storage device for storing one or more two-dimensional arrays of data elements that can be accessed by the processing circuitry of the data processing system when performing data processing operations. This can provide an efficient mechanism for performing several different types of operations, such as outer product operations. Considering two input operand vectors, the outer product of these two vectors is a matrix of data elements generated by multiplying each data element of one operand by each data element of the other operand. If the two vectors have dimensions M and N, their outer product is an M×N matrix. The provision of an array storage device that can store a two-dimensional array of data elements can provide a useful mechanism for storing the results of such outer product operations.
[0003] Outer product operations can be useful in modern data processing systems when implementing various types of calculations. For example, the use of outer product operations can be used to accelerate matrix multiplication. However, in order to improve efficiency and processing ability, it is desirable to use the array storage device efficiently and improve the utilization of the processing circuitry / multiply-accumulate resources provided within the data processing system.
Summary of the Invention
[0004] In one exemplary configuration, an apparatus is provided that includes a processing circuit that performs vector operations, an instruction decoder circuit that decodes instructions from an instruction set to control the processing circuit to perform the vector operations specified by the instructions, and an array storage device that includes storage elements for storing data elements and is configured to store at least one two-dimensional array of data elements accessible to the processing circuit when performing vector operations. The instruction set includes a multiple outer product instruction that identifies a first source vector operand, a second source vector operand, and a given two-dimensional array of data elements in the array storage device that form a destination operand. At least the first source vector operand identifies at least one vector of data elements treated as including a plurality of sub-vectors, at least the second source vector operand identifies a plurality of vectors of data elements, and the instruction decoder circuit is configured to control the processing circuit to perform an outer product operation for each sub-vector identified by the first source vector operand in response to the multiple outer product instruction. Each outer product operation includes multiplying each data element of each associated sub-vector identified by the first source vector operand by each data element of a group of data elements selected from the second source vector operand to generate a plurality of outer product results, and using each outer product result to update a value held in an associated storage element in a given two-dimensional array of storage elements. The processing circuit includes a selection circuit that controls the selection of data elements processed by each outer product operation to switch the vectors of the second source vector operand when switching between different sub-vectors within a given vector of the first source vector operand.
[0005] In another exemplary configuration, a method for performing an outer product operation is provided. This method includes using a processing circuit to perform a vector operation, using an instruction decoder circuit to decode an instruction from an instruction set to control the processing circuit to perform the vector operation specified by the instruction, and providing an array storage device comprising storage elements for storing data elements, the array storage device being configured to store at least one two-dimensional array of data elements accessible to the processing circuit when performing a vector operation. The instruction set includes a multiple outer product instruction that identifies a first source vector operand, a second source vector operand, and a given two-dimensional array of data elements in the array storage device that forms a destination operand. At least the first source vector operand identifies at least one vector of data elements treated as including a plurality of sub-vectors, and at least the second source vector operand identifies a plurality of vectors of data elements. Providing the array storage device, and in response to the multiple outer product instruction being decoded by the instruction decoder circuit, controlling the processing circuit to perform an outer product operation on each sub-vector identified by the first source vector operand, wherein each outer product operation includes multiplying each data element of each associated sub-vector identified by the first source vector operand by each data element of a group of data elements selected from the second source vector operand to generate a plurality of outer product results, and using each outer product result to update a value held in an associated storage element within a given two-dimensional array of storage elements. Controlling the processing circuit to include, and when switching between different sub-vectors within a given vector of the first source vector operand, controlling the selection of data elements processed by each outer product operation to switch between the vectors of the second source vector operand.
[0006] In a further exemplary configuration, a computer program is provided for controlling a host data processing device to provide an instruction execution environment. The computer program includes a processing program logic for performing vector operations, an instruction decoding program logic for decoding instructions from an instruction set to control the processing program logic to perform the vector operations specified by the instructions, and an array storage device emulation program logic for emulating an array storage device having storage elements for storing data elements. The array storage device is configured to store at least one two-dimensional array of data elements accessible to the processing program logic when performing vector operations. The instruction set includes a multiple outer product instruction for identifying a first source vector operand, a second source vector operand, and a given two-dimensional array of data elements in the array storage device that form a destination operand. At least the first source vector operand identifies at least one vector of data elements treated as including a plurality of sub-vectors, and at least the second source vector operand identifies a plurality of vectors of data elements. The instruction decoding program logic is configured to control the processing program logic to perform an outer product operation for each sub-vector identified by the first source vector operand in response to the multiple outer product instruction. Each outer product operation includes multiplying each data element of each associated sub-vector identified by the first source vector operand by each data element of a group of data elements selected from the second source vector operand to generate a plurality of outer product results, and updating the values held in the associated storage elements in a given two-dimensional array of storage elements using each outer product result. The processing program logic includes a selection program logic for controlling the selection of data elements processed by each outer product operation to switch the vectors of the second source vector operand when switching between different sub-vectors within a given vector of the first source vector operand.
Brief Description of the Drawings
[0007] This technology will be further described by way of example only, with reference to the embodiments of this technology shown in the accompanying drawings.
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12A
Figure 12B
Figure 13
Figure 14
Figure 15
Modes for Carrying Out the Invention
[0008] In an exemplary implementation considered herein, there is provided an apparatus having a processing circuit for performing vector operations and an instruction decoder circuit for decoding those instructions from an instruction set to control the processing circuit to perform the vector operations specified by the instructions. An array storage device is also provided that includes storage elements for storing data elements. The array storage device is configured to store at least one two-dimensional array of data elements accessible to the processing circuit when performing vector operations.
[0009] As described above, the use of an array storage device can provide a useful mechanism for performing certain operations, such as outer product operations. In particular, the matrix of data elements generated as a result of performing an outer product operation can be stored within the associated data elements of the two-dimensional array within the array storage device. However, efficiently using the available storage elements within a given two-dimensional array can be beneficial as it can help improve the processing capabilities of the system. It is also beneficial to better utilize the hardware multiply-accumulate resources provided within the system, and in some implementations, for each storage element within the array, there may already be provided within the system to support other calculations that can be performed using the two-dimensional array within the array storage device.
[0010] According to the technology described in this specification, the instruction set is configured to include a "multiple outer product instruction" such as an instruction that identifies two source vector operands, and a given two-dimensional array of data elements in an array storage device that forms a destination operand. Further, at least one of the source vector operands (referred to herein as the "first source vector operand", but note that this first source vector operand can be either of the two source vector operands specified by the multiple outer product instruction and does not have to be the first input operand specified by the instruction) identifies at least one vector of data elements that is treated as including a plurality of sub-vectors. Further, at least the other of the source vector operands (referred to herein as the "second source vector operand", but note that this second source vector operand can be either of the two source vector operands specified by the multiple outer product instruction and does not have to be the second input operand specified by the instruction) identifies a plurality of vectors of data elements.
[0011] The command decoder is configured to control the processing circuit to perform an outer product operation for each sub-vector identified by the first source vector operand in response to a multiple outer product command. Each outer product operation includes multiplying each data element of a group of data elements selected from the second source vector operand by each data element of an associated sub-vector identified by the first source vector operand to generate a plurality of outer product results, and using each outer product result to update a value held in an associated memory element within a given two-dimensional array of memory elements. The group of data elements selected from the second source vector operand may depend on whether the second source vector operand is also considered to include a plurality of sub-vectors. Thus, in an exemplary implementation, when the second source vector operand is considered to include a plurality of sub-vectors, the group of data elements selected from the second source vector operand may be data elements belonging to a selected sub-vector of the second source vector operand, but when the second source vector operand is not considered to include a plurality of sub-vectors, the group of data elements selected from the second source vector operand may be data elements belonging to a selected vector of the second source vector operand.
[0012] The processing circuit further includes a selection circuit that controls the selection of data elements processed by each outer product operation to switch vectors of the second source vector operand when switching between different sub-vectors within a given vector of the first source vector operand.
[0013] In some exemplary use cases, the inventors have recognized that when performing an outer product operation using two source vectors, one or both dimensions of those source vectors may be smaller than the corresponding dimensions of the two-dimensional array used to store the result of the outer product operation. This can lead to inefficient use of the memory elements of the two-dimensional array, as a significant number of memory elements may not be used, and can also lead to inefficient use of the resources of the hardware components that form the processing circuit (which may be capable of performing calculations to generate results for each of the memory elements). However, according to the techniques described herein, it is possible to define a single instruction (i.e., the multiple outer product instruction described above) that enables multiple outer product operations to be performed through the use of sub-vectors within one or both of the source vector operands, and the result of each outer product operation is stored within the associated memory elements of a two-dimensional (2D) array. This can significantly improve throughput by enabling multiple outer product operations to be performed in response to a single instruction (in an exemplary implementation, those multiple outer product operations may be performed in parallel), while also making more efficient use of the available memory elements within the array memory device.
[0014] Through the execution of the multiple outer product instruction described herein, one or more rows and / or columns of the 2D array can be configured to capture the results generated for one or more outer product operations by using two or more sub-vectors provided within a given input vector when calculating the outer product results used to update those one or more rows and / or columns.
[0015] In other words, for each vector register within a single source vector operand, a separate outer product operation is performed, and thus the total number of outer products is determined by the product of the number of vectors provided by both source vector operands in an exemplary implementation. Further, each outer product operation uses a subset (sub-vector) of the input data elements of at least one source vector operand.
[0016] Within any given vector of data elements that is treated as containing a plurality of subvectors, the various subvectors can be considered to occupy associated subvector regions within the given vector of data elements. In some cases, each subvector may occupy the entire associated subvector region, but in other exemplary implementations, some data element positions within a given vector may not be used, and thus one or more of the subvectors may not occupy the entire associated subvector region. Any particular subvector region need not be provided by adjacent data element positions within a given vector, and thus it should also be noted that the various data elements forming any particular subvector need not be provided adjacent to each other within the given vector.
[0017] The selection circuit can take various forms, but in one exemplary implementation, the selection circuit may include an associated multiplexer circuit for selecting data elements from a first source vector operand and a second source vector operand provided to a multiplication operation for each multiplication operation performed by the processing circuit to generate a corresponding outer product result, and the selection performed by the associated multiplexer circuit is controlled according to which outer product operation the corresponding outer product result pertains to. If the hardware implementation provides a separate multiplier circuit for each of the multiplication operations being performed, such an approach may enable the various multiplication operations to be performed in parallel.
[0018] At least a first source vector operand identifies a plurality of sub-vectors (provided by one or more vectors), and at least a second source vector operand identifies a plurality of vectors (as described above, either one of the source vector operands specified by the instruction can be considered as the first source vector operand, and thus the other source vector operand is considered as the second source vector operand), but both the first source vector operand and the second source vector operand are capable of identifying a plurality of sub-vectors, and in fact, both the first source vector operand and the second source vector operand are also capable of identifying a plurality of vectors. In a particular exemplary implementation, both the first source vector operand and the second source vector operand identify a plurality of vectors of data elements, and each vector is treated as including a plurality of sub-vectors. In such a configuration, it is possible to execute at least four outer product operations in response to a single multiple outer product instruction. In such an implementation, the number of sub-vectors specified by each source vector operand can in principle be different, but in a particular exemplary configuration, each of the first source vector operand and the second source vector operand includes N sub-vectors, and the processing circuit is configured to execute N outer product operations.
[0019] In an exemplary implementation where both the first source vector operand and the second source vector operand identify a plurality of vectors of data elements, and each vector includes a plurality of sub-vectors, the plurality of sub-vectors within each vector of the first source vector operand can be considered to have sub-vectors associated within different vectors of the second source vector operand. Next, the selection circuit is configured to control the selection of data elements processed by each outer product operation such that when switching different sub-vectors within a given vector of the first source vector operand, data elements from the associated sub-vectors within the second source vector operand are selected, and a switch to a different vector of the second source vector operand is made.
[0020] In an alternative exemplary implementation, only one of the source vector operands may identify sub-vectors. For example, the first source vector operand may include P plural sub-vectors, the second source vector operand may include P vectors of data elements, and each vector within the second source vector operand is associated with one of the sub-vectors within the first source vector operand. Such an exemplary implementation still allows a plurality of outer product operations to be executed in response to a single instruction and enables efficient utilization of available memory elements within a two-dimensional array.
[0021] In such an exemplary configuration, the processing circuit may be configured to execute P outer product operations, and each outer product operation is executed using the associated sub-vector from the first source vector operand and the associated vector from the second source vector operand as inputs.
[0022] There are various ways to specify sub-vectors within a given vector of source vector operands. In an exemplary implementation, for at least one given vector that is treated as including multiple sub-vectors, the data elements forming each sub-vector are provided at adjacent data element positions within the given vector.
[0023] However, it is not a necessary condition that the data elements forming the sub-vectors be provided at adjacent data element positions. In fact, for at least one given vector that is treated as including multiple sub-vectors, the data elements forming each sub-vector may be provided at non-adjacent data element positions within the given vector. Such an approach allows for greater flexibility as to how sub-vectors are arranged within a particular vector, for example, allowing data elements of one sub-vector to be interleaved with data elements of another sub-vector. In some implementations, it may be possible to have one or more vectors with interleaved sub-vector elements and one or more vectors with adjacent sub-vector elements, but in an exemplary implementation, the same scheme is used for each of the vectors that include sub-vectors.
[0024] In some implementations, any given vector that is considered to be formed from a plurality of sub-vectors is configured such that all data element positions contain a valid data element of one of the sub-vectors. However, this is not a requirement, and in alternative implementations, for at least one given vector that is treated as including a plurality of sub-vectors, the given vector may have one or more unused data element positions that do not contain data elements of the plurality of sub-vectors. This may result in some of the memory elements in the destination 2D array not being used, but still enables the execution of a plurality of outer product operations in response to a single instruction, thereby enabling an improvement in processing power / throughput, and also enabling an improvement in the utilization of the 2D array as compared to existing schemes that execute only a single outer product operation in response to a single instruction. There are several ways to identify unused data element positions, but in an exemplary implementation, for example, a predicate technique is used by specifying a predicate vector operand in relation to one or more of the source vector operands, and such a predicate vector operand provides one or more vectors of predicate values to identify for each data element position whether the data element at that position is included in the outer product operation.
[0025] There are various ways to specify the source vector operands of the multiple outer product instruction. However, in an exemplary implementation, the apparatus further comprises a set of vector registers accessible to the processing circuit, each vector register being configured to store a vector containing a plurality of data elements, and the first source vector operand and the second source vector operand include vectors contained within vector registers of the set of vector registers. Thus, the multiple outer product instruction can identify each source vector operand by specifying one or more vector registers within the set of vector registers whose contents are to form that source vector operand.
[0026] In an exemplary implementation, the vector length identifies the size of the vector registers within a set of vector registers and the size of a given two-dimensional array of data elements within the array storage device (e.g., using the vector length, both the x-dimension and y-dimension of the two-dimensional array can be specified). Some architectures may support variable vector lengths, and for any particular instantiation of the device, the vector length may be fixed, but the vector length may vary between different instantiations of the device, and the same instructions can be executed on any of those different instantiations of the device. Often, the outer product operations performed use one or more vectors / sub-vectors that are shorter than the specified vector length, and thus there is a high likelihood of opportunities to use the techniques described herein to improve processing power and 2D array utilization, so the techniques described herein can be particularly beneficially used within devices with relatively long vector lengths.
[0027] In an exemplary implementation, the multiple outer product instruction may be configured to provide a sub-vector indicator used to determine the number of sub-vectors within each vector that is treated as including multiple sub-vectors, and the size of each sub-vector depends on the determined number of sub-vectors and the vector length. In a particular exemplary implementation, the sub-vector indicator may be specified in a way that does not depend on the vector length. For example, the sub-vector indicator may be configured to identify that the vector is divided into two sub-vector regions (e.g., by identifying that each sub-vector region is half of the vector length), four sub-vector regions (e.g., by identifying that each sub-vector region is a quarter of the vector length), etc., and thus the actual size of each sub-vector region depends on the vector length.
[0028] In some implementations, each sub-vector may occupy the entire associated sub-vector region, but this is not a requirement. In alternative implementations, a sub-vector may occupy only a portion of the associated sub-vector region, and the remaining portion may be unused (i.e., include one or more unused data element positions). In such cases, there are various ways in which the unused data element positions can be handled. For example, the hardware may use the data elements within all of the data element positions to compute a result, and then simply ignore the results that are not needed later (i.e., effectively process the input as if there were no unused data element positions). However, alternatively, the aforementioned prediction techniques can be used to identify individual data element positions that should not be used when performing the outer product calculation. The use of such prediction techniques can avoid generating unnecessary results, and thus more easily facilitate the merging of the valid result data elements with the existing contents of the 2D array.
[0029] The sub-vector indicator can be specified in various ways. For example, the sub-vector indicator may be an explicit field provided within the instruction, or alternatively, may be implicitly specified by being part of the opcode used to define the instruction. Thus, a particular form of the multiple outer product instruction may have an explicit sub-vector indicator field that allows the number of sub-vectors to be specified, or alternatively, there may be different variants of the multiple outer product instruction for each of the different numbers of sub-vectors supported. In an exemplary implementation where an explicit sub-vector indicator field is provided, this may, for example, enable the sub-vector indicator to be set at runtime by providing, within the sub-vector indicator, the register identifier of a register whose contents define the number of sub-vectors.
[0030] The outer product operation executed in response to a multiple outer product command can take various forms. However, in an exemplary implementation, the multiple outer product command is an accumulation command, and each outer product result is used to update an existing value held in an associated memory element within a given two-dimensional array of memory elements by combining the outer product result with the existing value. The way the outer product result is combined with the existing value can vary depending on the implementation, but in an exemplary implementation, it can involve either adding the outer product result to the existing value or subtracting the outer product result from the existing value. The use of a two-dimensional array can be particularly beneficial when the outer product operation can be executed repeatedly multiple times, each of which executes an accumulation operation that generates a result to be accumulated within the two-dimensional array.
[0031] In an exemplary configuration, there is a one-to-one correspondence between each generated outer product result and its corresponding associated memory element within the two-dimensional array, but this is not a requirement. In other exemplary implementations, the outer product operation executed may be such that multiple generated outer product results are associated with the same memory element within the two-dimensional array. As a specific example, the multiple outer product command may be a sum of outer products command, and as a result, multiple outer product results have the same associated memory element within a given two-dimensional array of memory elements, and those multiple outer product results are combined to update the value held in the associated memory element. For example, each of multiple outer product results associated with the same memory element may be added together when updating the value held in the associated memory element. Similar to the previous embodiments, an accumulation variant may be supported, whereby the sum of the results of various outer product results associated with a particular memory element is added to or subtracted from the current value within the associated memory element to generate a new value to be stored within that memory element. When performing such a sum of outer products operation, typically this is the case when the individual data elements provided within the source vector operand are smaller than the data element size associated with each memory element within the two-dimensional array.
[0032] The techniques described herein provide a great deal of flexibility in how source vector operands are specified, including how many sub-vectors within various source vector operands are specified, and enable the improvement of the processing capabilities and utilization of 2D arrays in many different scenarios. As a mere specific example, in one particular use case, both the first source vector operand and the second source vector operand each contain two vectors, each vector being formed from two sub-vectors, and the instruction decoder circuit is configured to control the processing circuit to perform four outer product operations in response to a multiple outer product instruction, and the results of these four outer product operations are stored in the memory elements within a relevant region of a given two-dimensional array of memory elements. In an implementation where the sub-vectors completely occupy each vector, this may enable the entire 2D array to be used to store the results of the four outer product operations.
[0033] Next, with reference to the figures, a particular exemplary implementation will be considered.
[0034] FIG. 1 schematically shows a data processing system 10 comprising a processor 20 coupled to a memory 30 that stores data values 32 and program instructions 34. The processor 20 includes an instruction fetch unit 40 for fetching program instructions 34 from the memory 30 and supplying the fetched program instructions to an instruction decoder circuit 50. The decoder circuit 50 decodes the fetched program instructions and generates control signals for controlling the processing circuit 60 to perform processing operations on data values held within the memory elements of the register storage device 65 as specified by the decoded vector instructions. As shown in FIG. 1, the register storage device 65 may be formed from a plurality of different blocks. For example, a scalar register file 70 may be provided that includes a plurality of scalar registers that may be specified by instructions, and similarly, a vector register file 80 may be provided that includes a plurality of vector registers that may be specified by instructions.
[0035] Also, as shown in FIG. 1, the processor 20 can access the array storage device 90. In the embodiment shown in FIG. 1, the array storage device 90 is provided as part of the processor 20, but this is not a requirement. In various embodiments, the array storage device can be implemented as any one or more of an architecturally addressable register, an architecturally non-addressable register, a scratchpad memory, and a cache.
[0036] In an exemplary implementation, the processing circuit 60 may include both a vector processing circuit and a scalar processing circuit. The general distinction between scalar processing and vector processing is as follows. Vector processing may involve applying a single vector processing instruction to the data elements of a data vector having a plurality of data elements at each position within the data vector. The processing circuit may also perform vector processing to execute operations on a plurality of vectors within a two-dimensional array (sometimes also referred to as a subarray) of data elements stored in the array storage device 90. Scalar processing effectively performs operations on a single data element rather than a data vector. Vector processing can be useful when processing operations are to be performed on many different instances of the data being processed. In a vector processing configuration, a single instruction can be applied simultaneously to a plurality of data elements (of a data vector). This can improve the efficiency and throughput of data processing compared to scalar processing.
[0037] The processor 20 can be configured to process a two-dimensional array of data elements stored in the array storage device 90. In at least some embodiments, the two-dimensional array can be accessed as a one-dimensional vector of data elements in a plurality of directions. In an exemplary implementation, the array storage device 90 may be configured to store one or more two-dimensional arrays of data elements, and each two-dimensional array of data elements may form a square array portion of a larger or higher-dimensional array of data elements in memory.
[0038] FIG. 2 shows an example of the architectural registers 65 of a processor 20 that may be provided in an exemplary implementation. Architectural registers (such as those defined in an Instruction Set Architecture (ISA)) may include a set of scalar integer registers 100 that serve as general-purpose registers for processing operations executed by the scalar processing circuitry within the processing circuitry 60. For example, a certain number of general-purpose registers 100 may be provided, such as 31 registers X0 to X30 in this example (the 32nd encoding of the scalar register field may not correspond to a register provided in hardware, which may, for example, be considered to indicate a value of 0 by default or may be used to indicate a dedicated type of register rather than a general-purpose register). It may be possible to access scalar registers of different sizes mapped to the same physical storage device. For example, the register labels X0 to X30 may refer to 64-bit registers, but the same registers may be accessed as 32-bit registers (e.g., accessed using the lower 32 bits of each 64-bit register provided in hardware), in which case the register labels W0 to W30 may be used in assembly code to refer to the same registers.
[0039] Also, the architectural registers available for selection by program instructions within the ISA supported by the decoder 50 may include a certain number of vector registers 105 (labeled Z0 to Z31 in this example). Of course, it is not essential to provide the number of scalar / vector registers shown in FIG. 2, and other embodiments may provide different numbers of registers that can be specified by program instructions. Each vector register may store a vector operand that includes a variable number of data elements where each data element may represent an independent data value. In response to vector processing (SIMD) instructions, the processing circuit may perform vector processing on the vector operands stored in the registers to generate a result. For example, the vector processing may include per-lane operations where corresponding operations are performed on each lane of elements within one or more operand vectors to generate corresponding results for the elements of the result vector. When performing vector or SIMD processing, each vector register may have a certain vector length VL, where the vector length refers to the number of bits within a given vector register. The vector length VL used in vector processing mode may be fixed for a given hardware implementation or may be variable. The ISA supported by the processor 20 may support a variable vector length so that different processor implementations can choose to implement vector registers of different sizes, but the ISA may not be dependent on the vector length such that the instructions are designed so that the code can function correctly regardless of the specific vector length implemented on a given CPU on which the program is executed.
[0040] The vector registers Z0 to Z31 may also function as operand registers for storing vector operands that provide inputs to the processing and accumulation operations performed by the processing circuit 60 on the two-dimensional array of data elements stored in the array storage device 90. When the vector registers are used to provide inputs to such operations, the vector registers may have the same vector length VL as used for vector operations or may have a vector length MVL that can be different.
[0041] As shown in FIG. 2, the architecture register includes a certain number N A of array registers 110, ZA0 to ZA(N A -1). Each array register can be regarded as a single 2D array of data elements, for example, a set of register storage devices for storing the results of processing and accumulation operations. However, the processing and accumulation operations may not be the only operations that can use the array register. The array register can also be used to store a square array when performing a row / column transposition of the array structure in the memory. When a program instruction refers to one of the array registers 110, it is referred to as a single entity using the array identifier ZAi, but some types of instructions (e.g., data transfer instructions) can also select a sub-part of the array by defining an index value that selects a part of the array (e.g., one horizontal / vertical group of elements).
[0042] In fact, the physical implementation form of the register storage device corresponding to the array register is also shown in FIG. 2, and it is a certain number N R of array vector registers ZAR0 to ZAR(N R-1) may be included. The array vector register ZAR forming the array register memory device 110 may be a register set separate from the vector registers Z0 to Z31 used for vector input to SIMD processing and array processing. Each of the array vector registers ZAR may have a vector length MVL, and thus each array vector register ZAR may store a 1D vector of length MVL that can be logically divided into a variable number of data elements. For example, when MVL is 512 bits, this can be, for example, a set of 64 8-bit elements, 32 16-bit elements, 16 32-bit elements, 8 64-bit elements, or 4 128-bit elements. It will be understood that not all of these options need to be supported in a given implementation. By supporting variable element sizes, this provides flexibility for computational processing involving data structures of different precisions. To represent a 2D array of data, the vector registers ZAR0 to ZAR(N R -1) group can be logically regarded as a single entity assigned a given one of the array register identifiers ZA0 to ZA(N A -1), and thus the 2D array is formed of elements that are expanded within a single vector register corresponding to one dimension of the array and elements of the other dimension of the array that are striped across multiple vector registers.
[0043] It is not essential but may be useful to configure the array register ZA to store a square array of data where the number of elements in the horizontal direction is equal to the number of elements in the vertical direction. This can be useful for supporting on-the-fly transposition of the array to switch the row / column dimension of the array structure in memory when transferring the array structure between the array register 110 and memory by providing support for reading from / writing to the array register 110 in either the horizontal or vertical direction. By providing support for writing to / reading data from the 2D array register in either the horizontal or vertical direction, this can enable data loaded from memory in one direction (e.g., row by row) to be written back to memory in the opposite direction (e.g., column by column) faster than is possible using some gather / scatter load / store or permutation operations for transferring data between memory and the vector register.
[0044] As described above, the processing circuit 60 is configured to access the scalar register 70, the vector register 80, and / or the array storage device 90 under the control of the instructions decoded by the decoder circuit 50. Here, further details of this latter configuration will be described with reference to FIG. 3A, which merely provides one exemplary embodiment of how the array storage device can be accessed, particularly considering access to the square 2D array within the array storage device.
[0045] In the illustrated embodiment, the square 2D array within the array storage device 90 is configured as an array 205 of n×n memory elements / locations 200, where n is an integer greater than 1. In this embodiment, n is 16, which means that the granularity of access to the memory locations 200 is 1 / 16 of the entire storage device in either the horizontal array direction or the vertical array direction.
[0046] From the perspective of the processing circuit, an n×n array of positions can be accessed as n linear (one-dimensional) vectors in a first direction (e.g., the horizontal direction as depicted) and n linear vectors in a second array direction (e.g., the vertical direction as depicted). Thus, the n×n memory locations are configured as, or at least accessible as, 2n linear vectors for each of the n data elements, as seen from the processing circuit 60.
[0047] The array of memory locations 200 is accessible by access circuits 210, 220, column selection circuit 230, and row selection circuit 240 under the control of a control circuit 250 that communicates at least with the processing circuit 60 and optionally with the decoder circuit 50.
[0048] Referring to FIG. 3B, the n linear vectors in the first direction (horizontal or “H” direction as depicted) are, in the case of an exemplary square 2D array designated as “A1” (note that there may be two or more such 2D arrays, e.g., A0, A1, A2, etc., provided within the array memory device 90 as will be discussed below), each of the 16 data elements 0 - F (in hexadecimal), and in this example, may be referred to as A1H0 - A1H15. The same underlying data stored in the 256 entries (16×16 entries) of the array memory device 90A1 of FIG. 3B may instead be referred to as A1V0 - A1V15 in the second direction (vertical or “V” direction as depicted). For example, note that data element 260 is item F of A1H0 but is referred to as item 0 of A1V15. Note that the use of “H” and “V” does not imply any spatial or physical layout requirements regarding the storage of the data elements that make up the array memory device 90, and in any exemplary application, is independent of whether the 2D array within the array memory device stores row data or column data.
[0049] FIG. 4 is a block diagram of an apparatus according to an exemplary implementation showing how a processing circuit is used to perform an outer product operation. A vector register file 80 provides a plurality of vector registers that can be used to store vectors of data elements. As described above, a multiple outer product instruction may be configured to identify a first source vector operand 300 and a second source vector operand 320. At least the first source vector operand 300 is configured to identify at least one vector of data elements 305 that is treated as including a plurality of subvectors, and at least the second source vector operand 320 is configured to identify a plurality of vectors of data elements 325, 330. The terms “first” and “second” used herein to refer to the two source vector operands are used solely as labels to distinguish the two source vector operands and do not imply any particular ordering as to how those operands are specified by an instruction. Thus, either of the source operand fields of the instruction may be used to specify the first source vector operand referenced above, and then the other of the source operand fields may be used to specify the second source vector operand referenced above.
[0050] Furthermore, according to the techniques described herein, although at least one of the two source vector operands may identify a plurality of subvectors and the other source vector operand may not, in some exemplary implementations, both source vector operands may identify a plurality of subvectors. Similarly, both source vector operands may specify a plurality of vectors, and thus the first source vector operand 300 may include not only the first vector 305 but also the second vector 310. Further, the number of vectors specified by any particular source vector operand is not limited to either one vector or two vectors, and in fact, more vectors may be specified by a particular source vector operand (e.g., four vectors or eight vectors).
[0051] The processing circuit 60 is controlled in response to a control signal received from the decoder circuit 50, and when the decoder circuit 50 decodes the aforementioned multiple outer product instruction, the processing circuit is controlled to perform an outer product operation on each sub-vector identified by the first source vector operand. As part of this process, these control signals control a selection circuit 340 provided by the processing circuit 60 to select the appropriate data elements to be processed by each outer product operation. Each outer product operation multiplies each data element of the associated sub-vectors identified by the first source vector operand by a group of data elements (either sub-vector data elements or vector data elements) selected from the second source vector operand to generate a plurality of outer product results, and then uses each outer product result to update the value held in the associated memory element within a given two-dimensional array 380 of memory elements in the array storage device 90.
[0052] The selection circuit 340 can be configured in various ways. In an exemplary implementation, it includes a multiplexer circuit provided for each multiplier used to generate an outer product result from two input data elements, and the multiplexer circuit is used to select the appropriate two input data elements for each multiplier. The selection circuit controls the selection of the data elements to be processed by each outer product operation such that when switching between different sub-vectors within a given vector 305 or 310 of the first source vector operand, the vectors 325, 330 of the second source vector operand are switched.
[0053] Next, the selected input data element is transferred to the multiplication circuit 350, which, in an exemplary implementation as described above, may include a multiplication circuit for each outer product result to be generated. Each outer product result is generated by multiplying two input data elements provided to the corresponding multiplier within the multiplication circuit 350. The outer product result may be provided directly to the array update circuit 370 for use in updating the memory elements within the 2D array 380, and each outer product result has an associated memory element within the 2D array 380 and is used to update the value held in that associated memory element. However, the outer product operation being performed is often an accumulation operation, and each generated outer product result is combined with an existing value stored in an associated memory element of the 2D array 380 (e.g., by adding the outer product result to the existing value or subtracting the outer product result from the existing value) using an optional accumulation circuit 360. The multiplication circuit 350 and the optional accumulation circuit 360 are shown as separate blocks in FIG. 4, but in an exemplary implementation may be provided as a combined block formed from a multiply-accumulate circuit.
[0054] The array update circuit 370 is used to control access to the associated memory elements within the 2D array 380 to ensure that each value received by the array update circuit is used to update the associated memory element within the 2D array 380. In one embodiment, the array update circuit 370 may be implemented using the access components 210 - 250 described above with reference to FIG. 3A.
[0055] The outer product operation is usefully employed within a data processing system for a variety of different reasons, and thus the ability to execute multiple outer product operations in response to a single instruction can provide a significant improvement in processing power / throughput and can more efficiently utilize the available storage device resources provided by the two-dimensional array within the array storage device 90. As merely an example of how the outer product operation can be used, the outer product operation can be used to perform a matrix multiplication operation. Matrix multiplication may, for example, involve multiplying a first M×K matrix of data elements by a K×N matrix of data elements to produce an M×N matrix result of data elements. This operation can be decomposed into a plurality of outer product operations (more specifically, K outer product operations, where K may be referred to as the “depth”), and each outer product operation multiplies each data element of M vectors of data elements from the first matrix by each data element of N vectors of data elements from the second matrix to generate an M×N matrix of result data elements stored within the 2D array, involving executing a sequence of multiply-accumulate operations. The results of the plurality of outer product operations can be accumulated within the same 2D array to produce the M×N matrix generated by performing the aforementioned matrix multiplication.
[0056] FIG. 5A schematically shows selection functions 415, 420 executed by the selection circuit 340 of FIG. 4 for an exemplary use case where a first source vector operand specifies one vector 405 formed from two subvectors, a second source vector operand specifies two vectors 400, 410, and each of these vectors is considered a single vector of data elements (rather than being partitioned into subvectors).
[0057] In this exemplary implementation, upon execution of the multiple outer product instruction, two outer product operations are performed. The selection function 415 selects the data elements used in the first outer product operation, and the selection function 420 selects the data elements used in the second outer product operation. As shown in FIG. 5A, in the first outer product operation, the first sub-vector within the vector 400 of the first source vector operand is used in combination with the first vector 405 of the second source vector operand. Similarly, in the second outer product operation, the second sub-vector within the vector 400 of the first source vector operand is used in combination with the second vector 410 of the second source vector operand. Thus, it can be seen that the selection circuit switches between different vectors 405, 410 of the second source vector operand when switching between different sub-vectors within the first source vector operand 400.
[0058] FIG. 5B shows another embodiment where both source vector operands are considered to contain two vectors, each of which contains two sub-vectors. Thus, the first source vector operand contains two vectors 425, 430, and the second source vector operand contains two vectors 435, 440. In this case, four outer product operations are performed, each having an associated selection function 445, 450, 455, 460 executed by the selection circuit 340. As can be seen from the figure, the selection function 445 associated with the first outer product operation uses the first sub-vector within the first vector 425 of the first source vector operand in combination with the first sub-vector of the first vector 435 of the second source vector operand, and the selection function 450 associated with the second outer product operation uses the second sub-vector within the first vector 425 of the first source vector operand in combination with the third sub-vector provided by the second vector 440 of the second source vector operand. Similarly, the selection function 455 associated with the third outer product operation uses the third sub-vector provided by the second vector 430 of the first source vector operand in combination with the second sub-vector of the first vector 435 of the second source vector operand, and the selection function 460 associated with the fourth outer product operation uses the fourth sub-vector provided by the second vector 430 of the first source vector operand in combination with the fourth sub-vector provided by the second vector 440 of the second source vector operand. Here too, it can be seen that the selection circuit switches between different vectors of the second source vector operand when switching between different sub-vectors within a given vector (either the first vector 425 or the second vector 430) of the first source vector operand.
[0059] FIG. 6 is a diagram schematically showing a specific example of a plurality of outer product operations that can be executed in response to the execution of a multiple outer product instruction. In this case, both source vector operands specify two vectors 465, 470 and 475, 480 respectively, and each of these four vectors is considered to include two sub-vectors. This enables four outer product operations to be executed in response to a single instruction, and these outer product operations are schematically shown in FIG. 6 as problem numbers 1-4. As can be seen schematically by the illustrated 2D array 490, the result of each of these four outer product operations can be stored within a single 2D array. In particular, in this case, each outer product operation generates a 2×2 matrix that can be stored within the associated quarter of the 2D array 490. Thus, this enables the entire 2D array to be utilized, provides a particularly efficient utilization of the 2D array, and significantly increases processing power / throughput by avoiding the need to execute separate instructions to perform each of the outer product operations. In this exemplary use case, the multiple outer product instruction is referred to as an FMOPA4 (Floating-point Multiply Outer Product and Accumulate, a floating-point multiply outer product and accumulation that can execute up to four outer products) instruction.
[0060] FIG. 7 shows another exemplary use case where only one of the source vector operands identifies multiple sub-vectors. In particular, one source vector operand specifies a single vector 510 formed from two sub-vectors, and the other source vector operand identifies two vectors 500, 505, each of which is considered to contain a single vector of data elements. In this case, two outer product operations are performed, and again the entire 2D array 515 is utilized, and thus it can be seen that a particularly efficient and high-throughput approach for performing two outer product operations is provided. The multiple outer product instruction is also referred to here as the FMOPA4 instruction. This particular instruction identifies the array register ZA0 as the 2D array 515, and identifies that one of the source vector operands is formed by two vector registers z0 and z1, and the other source vector operand is formed by the vector register z4. The subscript ".s" indicates that the data being processed is single-precision floating-point data. In this particular embodiment, the instruction may also specify predicates for both the first source vector operand and the second source vector operand, enabling the outer product operation to be performed on selected data elements within the various vectors that form the two source vector operands while excluding other data elements. The subscript " / M" shown next to the predicate represents "merging", meaning that the current value stored in any storage element of the 2D array associated with the predicated (i.e., unused) data elements remains unmodified if the outer product operation is performed. This is in contrast to the subscript " / Z" which indicates that such values should be set to zero.
[0061] FIG. 8 shows a further example of a multiple outer product operation that can be performed in response to a multiple outer product instruction, where the instruction is referred to as an FMOPA16 (floating point multiply outer product and accumulate that can perform up to 16 outer products) instruction. In this case, one of the source vector operands specifies a single vector 528 consisting of four sub-vectors, and the other source vector operand specifies four vectors 520, 522, 524, 526, each of which is considered to contain a single vector of data elements. As can be seen from the figure, in this example four outer product operations are performed, each generating a 2×8 matrix, and all of the outer product results are accommodated within a single 2D array 530.
[0062] FIG. 9 shows a further exemplary scenario where each of the source vector operands is formed from two vectors 536, 538 and 532, 534 respectively, and each of the vectors is considered to contain two sub-vectors. In this example, the sub-vectors do not actually occupy the entire sub-vector region. In particular, in this exemplary use case, each sub-vector contains three data elements, but the associated sub-vector region provides four data element positions within the associated vector. Again, four outer product operations are performed, in this case each outer product operation generating a 3×3 matrix that can be stored within an associated region of a single 2D array 540.
[0063] FIG. 10 shows a further exemplary use case where each of the source vector operands is considered to include two vectors 546, 548 and 542, 544 respectively, and each vector includes two sub-vectors. However, in this case, the data elements forming each sub-vector are not arranged adjacent to each other within their associated vector. Instead, the data elements of different sub-vectors are interleaved. Supporting such flexibility can be useful as it can avoid the need to rearrange the data elements within the vector before performing the required outer product operation. Again, as in the embodiment of FIG. 9, it can be seen that each sub-vector includes three data elements and each outer product operation generates a 3×3 matrix that can be accommodated within a single 2D array 550. However, in this embodiment, the component data elements of each matrix are not stored in adjacent memory elements of the 2D array 550. Instead, the memory elements associated with the results of each outer product operation are separated from each other within the 2D array.
[0064] In the embodiment discussed with reference to FIGS. 6-10, each outer product result generated when performing each outer product operation has an associated memory element, and only one outer product result is used to update each associated memory element. However, the techniques described herein may also be used when performing other types of outer product operations, such as outer product sum operations. Such a use case is shown in FIG. 11, where a multiple outer product instruction (in this case, a multiple outer product sum instruction) is used to perform four outer product sum operations. As can be seen from the figure, each of the two source vector operands specifies two vectors 556, 558 and 552, 554, respectively, each of which is considered to be formed from two subvectors. However, the data element size within each subvector is reduced, and two outer product results are associated with each memory element within the 2D array 560. Thus, as a specific example, both the first outer product result calculated by multiplying a0 by i0 and the second outer product result calculated by multiplying b0 by j0 are added to generate a value used to update the associated memory element 562. In this embodiment, it can be seen that four outer product sum operations (eight raw outer product operations) are performed, each generating a 4×4 matrix, and the results of each of those outer product operations can all be accommodated within a single 2D array 560. In this embodiment, the instruction is referred to as a BFMOPA4 instruction, and the "B" in "BFMOPA4" indicates the BFloat16 data type.
[0065] FIG. 12A shows how an outer product result can be associated with a specific memory element within a 2D array. In the embodiment of FIG. 12A, a data element 570 from the first source vector operand is multiplied by a data element 572 from the second source vector operand using a multiplication function 574 to generate an outer product result, which is then subjected to an accumulation operation by an accumulation function 576 to generate an updated value stored in the associated memory element, either by adding the outer product result to the current value stored in the associated memory element 578 (or subtracting the outer product result from the current value stored in the associated memory element 578).
[0066] FIG. 12B shows an outer product sum operation in which two outer product results are associated with the same memory element in a 2D array. In this embodiment, to generate the first outer product result, the multiplication function 584 is used to multiply the data element 580 from the first source vector operand by the data element 582 from the second source vector operand. Similarly, the data element 586 from the first source vector operand and the data element 588 from the second source vector operand are multiplied by the multiplication function 590 to generate the second outer product result. Next, the two outer product results are added using the addition function 592, and the accumulation function 594 is executed to generate an updated data value for storage in the associated memory element 596. Thus, in some implementations, it will be understood that there can be two or more outer product results associated with the same memory element in a 2D array.
[0067] FIG. 13 is a diagram schematically showing fields that can be provided within a multiple outer product instruction according to one embodiment. The operation code field 605 is used to identify the type of instruction, in this case, to identify that the instruction is a multiple outer product instruction. Next, the subvector indicator field 610 can be used to identify the number of subvectors within each of one or more vectors considered to include subvectors. As described above, in one exemplary implementation, the subvector indication value within the field 610 can be specified in a manner independent of the vector length, for example, by specifying each subvector region to be a particular fraction of the vector length (thus, the actual size of each subvector region depends on the vector length, but knowledge of the vector length is not required when specifying the subvector indication value). Also as described above, in an alternative exemplary implementation, an explicit subvector indicator field 610 may not be provided, and instead, an indication of the subvector size may be directly encoded within the operation code field 605, effectively providing different variants of the instruction for different subvector sizes.
[0068] One or more other control information fields 615 may be provided, for example, to identify one or more predicates as mentioned above. Next, field 620 is used to identify one of the source vector operands, for example, by specifying one or more vector registers within vector register file 80 that provide the data elements of that source vector operand. Similarly, field 625 may again be used to identify the other source vector operand, for example, by specifying one or more vector registers within vector register file 80 that provide the data elements of that source vector operand. As described above, either one of fields 620, 625 can be used to specify the aforementioned first source vector operand, and the other field specifies the second source vector operand. Finally, field 630 can be used to identify the destination 2D array within array storage device 90 that is used to store the matrix generated as a result of performing the multiple outer product operations specified by the multiple outer product instruction.
[0069] FIG. 14 is a flowchart showing the steps performed during the decoding of a multiple outer product instruction according to an exemplary implementation. In step 650, it is determined whether a multiple outer product instruction has been encountered. If not, in step 655, standard decoding of the associated instruction is performed, and the processing circuit is controlled to execute the necessary operations defined by that instruction.
[0070] However, if a multiple outer product instruction is encountered, in step 660, the instruction is decoded to identify both source vector operands, the destination 2D array, the necessary subvector information, and the form of the outer product being performed (e.g., whether an accumulative outer product is being performed or a non-accumulative variant is being performed, and also for example, whether a normal outer product operation is being performed or an outer product sum operation is being performed).
[0071] Next, in step 665, the processing circuit is controlled to perform the necessary outer product operations and perform the necessary updates to the 2D array memory elements. As part of this process, the selection circuit is controlled to select data elements for each multiplication operation according to the identified subvectors. As described above, this involves switching between different subvectors within any given vector of the first source vector operand, which involves switching between vectors of the second source vector operand.
[0072] FIG. 15 shows an implementable form of a simulator. The foregoing examples implement the present invention from the perspective of apparatuses and methods for operating specific processing hardware that supports the technology, but it is also possible to provide an instruction execution environment according to the embodiments described herein, and the instruction execution environment is implemented by the use of a computer program. Such a computer program is often referred to as a simulator as long as the computer program provides a software-based implementation of a hardware architecture. Various simulator computer programs include binary translators including emulators, virtual machines, models, and dynamic binary translators. Typically, the simulator implementation form may be executed on a host processor 715, and optionally execute a host operating system 710 to support a simulator program 705. In some configurations, there may be multiple layers of simulation between the hardware and the provided instruction execution environment, and / or multiple different instruction execution environments may be provided on the same host processor. Historically, powerful processors have been required to provide simulator implementation forms that execute at a reasonable speed, but such an approach may be justified in certain situations, such as when it is desired to execute native code for another processor for reasons of compatibility or reuse. For example, the simulator implementation form may provide an instruction execution environment having additional functionality not supported by the host processor hardware, or may typically provide an instruction execution environment associated with a different hardware architecture. An overview of simulation is provided in "Some Efficient Architecture Simulation Techniques", Robert Bedichek, Winter 1990, USENIX Conference, Pages 53~63.
[0073] To the extent that embodiments have been previously described with reference to specific hardware constructs or features, in simulated implementations, equivalent functionality may be provided by suitable software constructs or features. For example, a particular circuit may be provided as computer program logic in a simulated implementation. Similarly, memory hardware such as registers or caches may be provided as software data structures in a simulated implementation. Also, the physical address space used to access memory 30 within hardware device 10 may be emulated as a simulated address space, which is mapped by simulator 705 to the virtual address space used by host operating system 710. In configurations where one or more of the hardware elements referred to in the foregoing examples are present in host hardware (e.g., host processor 715), some simulated implementations may use the host hardware, if suitable.
[0074] The simulator program 705 may be stored in a computer-readable storage medium (which may be a non-transitory medium), and provides a virtual hardware interface (instruction execution environment) to the target code 700 (which may include an application, an operating system, and a hypervisor). The virtual hardware interface is the same as the hardware interface of the hardware architecture modeled by the simulator program 705. Thus, the program instructions of the target code 700 may be executed from within the instruction execution environment using the simulator program 705. For this reason, a host computer 715 that does not actually have the hardware features of the device 10 described above can emulate these features. The simulator program may include processing program logic 720 that emulates the behavior of the processing circuit 60, instruction decode program logic 725 that emulates the behavior of the instruction decoder circuit 50, and array storage device emulation program logic 722 that maintains a data structure to emulate the array storage device 90. Thus, in the embodiment of FIG. 15, the technology described herein may be executed in software by the simulator program 705.
[0075] As is apparent from the foregoing discussion of exemplary implementations of the present technology, the technology described herein provides a single instruction (i.e., the multiple outer product instruction described above) that enables multiple outer product operations to be performed through the use of sub-vectors within one or both of the source vector operands. The result of each outer product operation is stored in an associated memory element of a selected two-dimensional array within the array storage device. This can significantly improve throughput by enabling multiple outer product operations to be performed in response to a single instruction, while also making more efficient use of the available memory elements within the array storage device.
[0076] In this application, the phrase "configured to" is used to mean that an element of an apparatus has a configuration capable of performing a defined operation. In this context, "configuration" means an arrangement or interconnection of hardware or software. For example, an apparatus may have dedicated hardware that provides a defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not mean that an element of an apparatus needs to be modified in any way to provide a defined operation.
[0077] Exemplary embodiments of the present invention have been described in detail herein with reference to the accompanying drawings, but it should be understood that the present invention is not limited to those exact embodiments, and various changes, additions, and modifications can be made by those skilled in the art without departing from the scope and spirit of the present invention as defined by the appended claims. For example, various combinations of the features of the dependent claims can be made with the features of the independent claims without departing from the scope of the present invention.
Claims
Claim 1 An apparatus, the apparatus comprising: A processing circuit that performs vector operations; An instruction decoder circuit that decodes the instruction from an instruction set to control the processing circuit to perform the vector operation specified by the instruction; An array storage device comprising storage elements for storing data elements, configured to store at least one two-dimensional array of data elements accessible to the processing circuit when performing the vector operation; and The instruction set includes a multiple outer product instruction that identifies a first source vector operand, a second source vector operand, and a given two-dimensional array of data elements in the array storage device that form a destination operand, wherein at least the first source vector operand identifies at least one vector of data elements treated as including a plurality of sub-vectors, and at least the second source vector operand identifies a plurality of vectors of data elements; The instruction decoder circuit is configured to control the processing circuit to perform an outer product operation for each sub-vector identified by the first source vector operand in response to the multiple outer product instruction, each outer product operation including multiplying each data element of an associated group of data elements selected from the second source vector operand by each data element of the associated sub-vector identified by the first source vector operand to generate a plurality of outer product results, and using each outer product result to update a value held in an associated storage element within a given two-dimensional array of the storage elements; The apparatus, wherein the processing circuit includes a selection circuit that controls the selection of the data elements processed by each outer product operation to switch the vectors of the second source vector operand when switching between different sub-vectors within a given vector of the first source vector operand. Claim 2 The selection circuit includes a multiplexer circuit associated with selecting data elements from the first source vector operand and the second source vector operand used in each multiplication operation executed by the processing circuit to generate a corresponding outer product result, and the selection performed by the associated multiplexer circuit is controlled according to which outer product operation the corresponding outer product result relates to. The apparatus according to claim 1.
3. The apparatus according to claim 1 or 2, wherein both the first source vector operand and the second source vector operand identify a plurality of vectors of data elements, and each vector is treated as including a plurality of sub-vectors.
4. The apparatus according to claim 3, wherein each of the first source vector operand and the second source vector operand includes N sub-vectors, and the processing circuit is configured to execute N outer product operations.
5. The plurality of sub-vectors within each vector of the first source vector operand have associated sub-vectors within different vectors of the second source vector operand. The selection circuit is configured to control the selection of the data elements processed by each outer product operation such that when switching between different sub-vectors within a given vector of the first source vector operand, data elements from the associated sub-vectors within the second source vector operand are selected, and switching to different vectors of the second source vector operand is performed. The apparatus according to claim 3 or 4.
6. The apparatus according to any one of claims 1 to 3, wherein the first source vector operand includes a plurality of P sub-vectors, the second source vector operand includes P vectors of data elements, and each vector within the second source vector operand is associated with one of the sub-vectors within the first source vector operand.
7. The processing circuit is configured to perform P outer product operations, each outer product operation being performed using an associated subvector from the first source vector operand and an associated vector from the second source vector operand as input values, the apparatus according to claim 6.
8. For at least one given vector treated as including a plurality of subvectors, the data elements forming each subvector are provided at adjacent data element positions within the given vector, the apparatus according to any one of claims 1 to 7.
9. For at least one given vector treated as including a plurality of subvectors, the data elements forming each subvector are provided at non-adjacent data element positions within the given vector, the apparatus according to any one of claims 1 to 8.
10. For at least one given vector treated as including a plurality of subvectors, the given vector may have one or more unused data element positions that do not include the data elements of the plurality of subvectors, the apparatus according to any one of claims 1 to 9.
11. A set of vector registers accessible to the processing circuit, each vector register being configured to store a vector including a plurality of data elements, the first source vector operand and the second source vector operand including vectors included within the vector registers of the set of vector registers, the apparatus further comprising a set of vector registers according to any one of claims 1 to 10.
12. The vector length identifies the size of the vector register within the set of vector registers and the size of a given two-dimensional array of the data elements within the array storage device, the multiple outer product instruction being configured to provide a subvector indicator used to determine the number of subvectors within each vector treated as including a plurality of subvectors, the size of each subvector depending on the determined number of subvectors and the vector length, the apparatus according to claim 11.
13. The subvector indicator is specified in a manner independent of the vector length, the apparatus according to claim 12.
14. The apparatus according to any one of claims 1 to 13, wherein the multiple outer product instruction is an accumulation instruction, and each outer product result is used to update the existing value held in the associated memory element in a given two-dimensional array of the memory element by combining the outer product result with the existing value.
15. The apparatus according to any one of claims 1 to 14, wherein the multiple outer product instruction is an outer product sum instruction, a plurality of outer product results have the same associated memory elements in a given two-dimensional array of the memory element, and their multiple outer product results are combined to update the value held in the associated memory element.
16. The apparatus according to any one of claims 1 to 15, wherein both the first source vector operand and the second source vector operand include two vectors, each vector is formed from two sub-vectors, and the instruction decoder circuit is configured to control the processing circuit to execute four outer product operations in response to the multiple outer product instruction, and the results of these four outer product operations are stored in the memory elements in an associated area of a given two-dimensional array of the memory element.
17. A method of performing an outer product operation, comprising: using a processing circuit that performs vector operations; using an instruction decoder circuit that decodes the instruction from an instruction set to control the processing circuit to perform the vector operation specified by the instruction; providing an array storage device comprising a memory element for storing data elements, the array storage device being configured to store at least one two-dimensional array of data elements accessible to the processing circuit when performing the vector operation; providing an array storage device, wherein the instruction set includes a multiple outer product instruction that identifies a first source vector operand, a second source vector operand, and a given two-dimensional array of data elements in the array storage device that form a destination operand, at least the first source vector operand identifies at least one vector of data elements treated as including a plurality of sub-vectors, and at least the second source vector operand identifies a plurality of vectors of data elements. In response to the multiple outer product command being decoded by the instruction decoder circuit, controlling the processing circuit to perform an outer product operation on each sub-vector identified by the first source vector operand, wherein each outer product operation multiplies each data element of each associated sub-vector identified by the first source vector operand by each data element of a group of data elements selected from the second source vector operand to generate a plurality of outer product results, and updating values held in associated memory elements within a given two-dimensional array of the memory element using each outer product result, and controlling the processing circuit as described above. A method including controlling the selection of the data elements processed by each outer product operation to switch the vectors of the second source vector operand when switching between different sub-vectors within a given vector of the first source vector operand. Claim 18 A computer program for controlling a host data processing device to provide an instruction execution environment, processing program logic for performing vector operations, instruction decoding program logic for decoding the instructions from an instruction set to control the processing program logic to perform the vector operations specified by the instructions, array memory device emulation program logic for emulating an array memory device comprising a memory element for storing data elements, the array memory device being configured to store at least one two-dimensional array of data elements accessible to the processing program logic when performing the vector operations, and including the array memory device emulation program logic as described above. The instruction set includes a multiple outer product instruction that identifies a first source vector operand, a second source vector operand, and a given two-dimensional array of data elements within the array memory device that forms a destination operand, at least the first source vector operand identifying at least one vector of data elements treated as including a plurality of sub-vectors, and at least the second source vector operand identifying a plurality of vectors of data elements. The command decoding program logic is configured to control the processing program logic to perform an outer product operation on each sub-vector identified by the first source vector operand in response to the multiple outer product command, and each outer product operation is configured to multiply each data element of the associated sub-vector identified by the first source vector operand by each data element of the group of data elements selected from the second source vector operand to generate a plurality of outer product results, and to update the values held in the associated memory elements within a given two-dimensional array of the memory element using each outer product result. A computer program including selection program logic that controls the selection of the data elements processed by each outer product operation so that when switching between different sub-vectors within a given vector of the first source vector operand, the vectors of the second source vector operand are switched.