Predicate Technology

JP2024525798A5Pending Publication Date: 2025-06-23ARM LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024502040
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-21
Filing Date
2022-06-22
Publication Date
2025-06-23

AI Technical Summary

Technical Problem

Existing vector processing technologies face inefficiencies in handling large amounts of predicate information, particularly in scalable vector processing where vector lengths and element sizes can vary, leading to impractical implementations of traditional predicate masking techniques.

Method used

The use of predicate counting techniques, where predicate data values are encoded with element size and element count, allowing efficient representation of consecutive identical predicate indicators, reducing the need for multiple predicate masks and facilitating scalable vector operations.

Benefits of technology

This approach allows for a more efficient encoding of predicate information, requiring fewer resources while supporting flexible and scalable vector processing operations, even in scenarios with varying vector lengths and element sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Apparatus, methods and programs are disclosed for predication of multiple vectors in vector processing. Encoding of predicate information including element size and element count is disclosed, where the predicate information includes multiple consecutive identical predicate indicators given by the element count, each predicate indicator corresponding to an element size.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present technology relates to data processing, and in particular to vector processing and the use of predicates to control the vector processing.

[0002] The data processing apparatus may comprise processing circuitry for performing vector processing operations, which may involve parallel operations performed on respective elements of a vector held in a vector register, the predicate of the vector processing operation on the vector including controlling which elements of the vector are subjected to the vector processing operation.

[0003] At least some examples described herein provide an apparatus, the apparatus comprising: a decode circuit for decoding instructions; a processing circuit for applying a vector processing operation specified by the instruction to an input data vector; The decode circuitry responds to the vector processing instruction by: Vector processing operations; one or more source operands; Given a source predicate register, generating control signals to cause the processing circuitry to perform vector processing operations on the source operands, the control signals further causing the processing circuitry to selectively apply the vector processing operations to elements of the one or more source operands predicated by a predicate indicator decoded from the predicate data value retrieved from the source predicate register; The predicate data value has an encoding, The element size and an element count indicating a number of consecutive identical predicate indicators, each predicate indicator corresponding to an element size;

[0004] At least some examples described herein provide a method of processing data, comprising: The data processing method is Decoding the instruction; controlling a processing circuit to apply a vector processing operation specified by the instruction to an input data vector; Decoding comprises, in response to the vector processing instruction: Vector processing operations; one or more source operands; Given a source predicate register, a control signal for causing the processing circuit to perform a vector processing operation on source operands, the control signal further comprising: generating a control signal for causing the processing circuit to selectively apply the vector processing operation to elements of the one or more source operands predicated by a predicate indicator decoded from a predicate data value retrieved from a source predicate register; The predicate data value has an encoding, The element size and an element count indicating a number of consecutive identical predicate indicators, each predicate indicator corresponding to an element size;

[0005] At least some examples described herein provide a computer program for controlling a host processing device to provide an instruction execution environment, the computer program comprising: decode logic for decode the instruction; processing logic for applying vector processing operations specified by the instructions to input data vectors; The decode logic performs the following operations in response to the vector processing instructions: Vector processing operations; one or more source operands; Given a source predicate register, generating control signals to cause the processing logic to perform vector processing operations on the source operands, the control signals further causing the processing circuitry to selectively apply the vector processing operations to elements of the one or more source operands predicated by a predicate indicator decoded from a predicate data value retrieved from the source predicate register; The predicate data value has an encoding, The element size and an element count indicating a number of consecutive identical predicate indicators, each predicate indicator corresponding to an element size; [Brief description of the drawings]

[0006] The present technology will now be further described, by way of example only, with reference to embodiments of the technology as illustrated in the accompanying drawings, to be read in conjunction with the following description. [Figure 1] 1 illustrates schematic diagrams of devices according to various configurations of the present technology; [Figure 2A] 1A-1C illustrate schematic diagrams of the use of vector processing instructions to specify vector processing operations that utilize predicate information encoded using a predicate-as-count approach, in accordance with various configurations of the present technology; [Figure 2B] 1A-1C illustrate schematic diagrams of the use of vector processing instructions to specify vector processing operations that utilize predicate information encoded using a predicate-as-count approach, in accordance with various configurations of the present technology; [Diagram 3] 1 illustrates generally the use of vector processing instructions to specify vector processing operations that utilize predicate information encoded using the predicate counting technique in accordance with various configurations of the present technology. [Figure 4A] 1 illustrates a schematic of the use of an inversion indicator in predicate encoding, according to various configurations of the present technique; [Figure 4B] 11A-11C illustrate schematic diagrams of varying positions of the boundary between element count indications and element size indications in the encoding used for predicate count definitions according to various configurations of the present technique; [Figure 4C] 10A-10C illustrate schematic diagrams of example bit values ​​that determine the location of the boundary between an element count indication and an element size indication in an encoding used for a predicate count definition, according to various configurations of the present technique. [Diagram 5] 1 illustrates a schematic of the encoding used for predicate count definition according to various configurations of the present technique; [Figure 6]10A-10C illustrate schematic diagrams of reading and writing predicate data values ​​that define a predicate count according to various configurations of the present technology; [Figure 7] 1 illustrates generally the use of predicate generating instructions according to various configurations of the present technology; [Figure 8] 1A-1C illustrate schematic diagrams of the use of a size indicator in a vector processing instruction that converts a predicate-as-counter into a predicate-as-mask, in accordance with various configurations of the present technology; [Figure 9A] 10A-10C illustrate schematic diagrams of the use of subpart indicators in predicate transformation instructions, according to various configurations of the present technology; [Figure 9B] 1 illustrates a schematic of an all-true predicate generation instruction according to various configurations of the present technique; [Figure 10] 1 illustrates a schematic of the use of a predicate count instruction according to various configurations of the present technology; [Figure 11] 1 illustrates a schematic representation of an implementation of a simulator that may be used.

[0007] One example disclosed herein includes an apparatus, the apparatus comprising: a decode circuit for decoding the instruction; a processing circuit for applying a vector processing operation specified by the instruction to an input data vector; The decode circuitry responds to the vector processing instruction by: Vector processing operations; one or more source operands; Given a source predicate register, generating control signals to cause the processing circuitry to perform vector processing operations on the source operands, the control signals further causing the processing circuitry to selectively apply the vector processing operations to elements of the one or more source operands predicated by a predicate indicator decoded from the predicate data value retrieved from the source predicate register; The predicate data value has an encoding, The element size and an element count indicating a number of consecutive identical predicate indicators, each predicate indicator corresponding to an element size;

[0008] The predication of a data processing operation is generally achieved by the use of a predicate mask, according to which the predicate mask holds a number of indicators corresponding to the possible parallel positions to which the data processing operation may be applied. The respective values ​​of the indicators determine whether the data processing operation at the respective positions should be performed or not. In this specification, this approach is referred to as "predicate masking". Thus, in the context of vector processing operations, a predicate mask may be provided that includes a number of indicators corresponding to the number of elements in a vector that are to undergo the vector processing operation. The respective values ​​of the indicators in the predicate mask then determine which elements of the vector are to undergo the vector processing operation. However, the inventors of the present technique have found that for some operations that require a large amount of predicate information, using multiple predicate masks to cover the data width is impractical to implement. For example, this may be the case in the context of multi-vector processing, and the problem may become even more severe in the context of scalable vector processing, i.e. when the data processing device is not constrained to perform vector processing on a fixed vector length with a fixed number of elements, but rather can perform such vector processing on scalable vectors, i.e. scalable vectors whose length and / or element size may vary.

[0009] In this context, the present technique provides an efficient way for a larger amount of predicate information to be provided by the contents of a source predicate register than is possible according to conventional predicate masking techniques. Instead of a one-to-one correspondence between the indicators (e.g., bit values) that form the contents of the source predicate register and the target elements of the operands of the instruction, as in the case of predicate masking, the present technique utilizes an encoding of the predicate data values ​​held in the source predicate register, the encoding specifying an element size and an element count. The element count indicates multiple consecutive identical predicate indicators, each predicate indicator corresponding to an element size. Although this technique, referred to herein as "predicate counting," does not support any individual setting of predicate indicators corresponding to individual elements of the source operand(s), it has been found that predicate usage often involves a set of active elements followed by a set of inactive elements (or vice versa) with no gaps in between. For example, when vector processing is dealing with matrix elements, a situation where a set of active elements is followed by inactive elements may arise at the end of a matrix row. Conversely, a situation where a set of inactive elements followed by active elements may arise at the beginning of a matrix row. If the encoding used by the present technique allows such a set of active / inactive elements to be efficiently encoded, it will require only a single source predicate register to provide the product data value, which can nevertheless represent the information needed for predicating multiple vector operands.

[0010] Thus, in some examples the required predicate indicator may comprise the full set of active elements, or conversely the full set of inactive elements, while more generally in other examples the processing circuitry is configured to decode the predicate data value to generate successive identical predicate indicators and further sequences of identical predicate indicators, where the successive identical predicate indicators and further sequences of identical predicate indicators comprise mutually inverse activity indications.

[0011] Nonetheless, in some cases where all predicate indicators may need to be the same (i.e., all active or all inactive), in some examples the processing circuitry generates all predicate indicators as consecutive identical predicate indicators in response to the element count having a predetermined value.

[0012] In some examples, the encoding of the predicate data value further comprises an inversion bit, and the repeat activity indication forming successive identical predicate indicators depends on the inversion bit. This inversion bit can thus be used to set the "polarity" of the mask information. Furthermore, the inversion bit thus defines the start of an active element (wherein this is the start of the set of elements if an active element is followed by an inactive element, or the point in the set of elements where an active element starts following the set of inactive elements if an inactive element is followed by an active element).

[0013] The encoding used may represent element size in a variety of ways and may be capable of representing a range of element sizes, but in some examples the encoding of the predicate data value uses an element size encoding to indicate the element size, and the element size encoding is The element size is a byte long, The element size is a halfword length, The element size is the word length, The element size is a doubleword length, The element size comprises an indication for at least one of being a quadword length.

[0014] The encoding of the predicate data value allows both element size and element count to be represented by the predicate data value, although the portions of the predicate data value representing these respective components need not be fixed. Indeed, in some examples, the encoding of the predicate data value includes a predetermined portion of the predicate data value that is used to indicate element size and element count, and the boundary location in the predetermined portion of the predicate data value between a first sub-portion indicating element size and a second sub-portion indicating element count depends on the indicated element size. This variable boundary between the two sub-portions allows flexibility in how the space available in the predicate data value is used. In particular, in examples where less space is needed to indicate element size, more space can be used for element count, and conversely, in examples where less space is needed to indicate element count, more space can be used for element size.

[0015] The particular manner in which the element count and element size are represented is not limited and may take a variety of forms, however, in some examples where the boundary positions are variable in the manner described above, the bit positions of the active bits in the first subportion indicate the element size and the bit positions of the active bits define the boundary positions.

[0016] An advantage of the present technique is the particularly efficient encoding used by the predicate counter, such that a large amount of predicate indicators can be represented by a relatively small amount of space in the predicate data value. Indeed, the device may be configured to process the predicate data value within a limited portion of the contents of the source predicate register. This may facilitate the ease of implementation of the present technique. For example, in some cases, the encoding of the predicate data value is limited to a predetermined number of bits of the predicate data value, and the processing circuitry is configured to ignore any further bits that may be held in the source predicate register beyond the bits that form the predetermined number of bits of the predicate data value when reading the predicate data value from the source predicate register. Similarly, in some cases, the encoding of the predicate data value is limited to a predetermined number of bits of the predicate data value, and the processing circuitry is configured to set any further bits that may be held in the target predicate register beyond the bits that form the predetermined number of bits of the predicate data value to a predetermined value when writing a new predicate data value to the target predicate register.

[0017] The present technology further proposes various additional instructions that a device may respond to to support efficient creation and use of predicate counter instances. Thus, in some instances, in response to a predicate generation instruction that specifies a predicate to be generated and a number of vectors to be controlled by the predicate to be generated, the decode circuitry generates a control signal that causes the processing circuitry to generate a predicate data value indicative of a corresponding element size and a corresponding element count.

[0018] In some examples, the decoding circuitry generates, in response to an all-true predicate generation instruction specifying an all-true predicate to be generated, a control signal that causes the processing circuitry to generate an all-true predicate data value indicative of all active elements for the predicate indicator.

[0019] In some examples, the decoding circuitry generates a control signal that causes the processing circuitry to generate an all-false predicate data value indicative of an all-false predicate data value indicative of an all-inactive element for the predicate indicator in response to an all-false predicate generation instruction that specifies an all-false predicate to be generated.

[0020] While the predicate counter representation provides a particularly efficient encoding density of predicate information, the present technique nevertheless recognizes that situations may arise in which a conventional predicate mask representation may be usefully employed, and therefore at least one instruction is proposed that enables conversion from a predicate counter representation to a predicate mask representation. Thus, in some examples, the decode circuitry, in response to a predicate conversion instruction that specifies a source predicate register that holds the predicate data value to be converted, generates control signals that cause the processing circuitry to decode the predicate data value to be converted and generate a converted predicate data value, the converted predicate data value comprising a direct mask style representation in which bit values ​​at bit positions indicate predicates of elements in the subject data item.

[0021] It is further proposed that if the predicate counter representation can easily cover a much larger number of elements than a comparably sized predicate mask representation, the predicate counter representation can be transformed into two or more predicate masks. Thus, in some examples, the predicate transformation instruction specifies two or more destination predicate registers, and the control signal causes the processing circuit to generate two or more transformed predicate data values, each of the two or more transformed predicate data values ​​including a direct mask style representation, and each of the two or more transformed predicate data values ​​corresponding to a different subset of the predicate indicators represented by the predicate data value to be transformed.

[0022] In some examples, the predicate transformation instruction specifies a multiplicity of two or more transformed predicate data values ​​to be generated.

[0023] In some examples, the predicate transformation instruction specifies which of multiple possible subsets of predicate bits represented by the predicate data value to be transformed is to be generated.

[0024] In some examples, the decode circuitry generates, in response to a predicate count instruction that specifies a source predicate register holding the predicate data values ​​to be counted, control signals that cause the processing circuitry to decode the predicate data values ​​to be transformed, determine a predicate indicator indicated by the predicate data values ​​to be transformed, and store a scalar value corresponding to the number of active elements in the predicate indicator in a destination general-purpose register.

[0025] In some examples, the predicate count instruction specifies an upper limit on the number of active elements to be counted, the upper limit corresponding to one of two vector lengths and four vector lengths.

[0026] The one or more source operands may indicate one or more data items of various types that are the subject of a vector processing operation. The type of data item is not a limitation of the present technology, so long as it includes multiple elements that may receive a predicate as part of the operation. Furthermore, there may be only one source operand, or there may be multiple source operands. In the case of multiple source operands, there is an implicit ordering of the operands, such that when a predicate is applied to all of those multiple source operands, a first portion of the predicate information encoded in the predicate data value is applied to the first source operand, a second portion of the predicate information encoded in the predicate data value is applied to the second source operand, and so on as appropriate. The source operands may indicate vector registers (in particular, they may be scalable vector registers), or they may indicate a contiguous range of memory locations. Indeed, the source vectors themselves may indicate a range of memory locations by pointers, i.e., by gather (load) or scatter (store) operations. In the case of a range of memory locations, the predicate controls which particular locations are / are not accessed for either the load or store operation.

[0027] Thus, in some examples, the one or more source operands comprise one source vector register, two source vector registers, or three source vector registers. Additional numbers of source vector registers are possible. In some examples, the one or more source operands include a range of memory locations. The range of memory locations may be a single contiguous block of memory locations or may be pointed to by a set pointer, and thus potentially distributed across a larger memory space.

[0028] In some examples, the vector processing instruction further specifies a destination vector register.

[0029] In some examples, the vector processing instruction further specifies a destination memory location.

[0030] One example disclosed herein is a method of processing data, the method comprising: Decoding the instruction; controlling a processing circuit to apply a vector processing operation specified by the instruction to an input data vector; Decoding comprises, in response to the vector processing instruction: Vector processing operations; one or more source operands; Given a source predicate register, a control signal for causing the processing circuit to perform a vector processing operation on source operands, the control signal further comprising: generating a control signal for causing the processing circuit to selectively apply the vector processing operation to elements of the one or more source operands predicated by a predicate indicator decoded from a predicate data value retrieved from a source predicate register; The predicate data value has an encoding, The element size and an element count indicating a number of consecutive identical predicate indicators, each predicate indicator corresponding to an element size;

[0031] In one example disclosed herein, there is a computer program for controlling a host processing device to provide an instruction execution environment, the computer program comprising: decode logic for decode the instruction; processing logic for applying vector processing operations specified by the instructions to input data vectors; The decode logic performs the following operations in response to the vector processing instructions: Vector processing operations; one or more source operands; Given a source predicate register, generating control signals to cause the processing logic to perform vector processing operations on the source operands, the control signals further causing the processing circuitry to selectively apply the vector processing operations to elements of the one or more source operands predicated by a predicate indicator decoded from a predicate data value retrieved from the source predicate register; The predicate data value has an encoding, The element size and an element count indicating a number of consecutive identical predicate indicators, each predicate indicator corresponding to an element size;

[0032] Some specific embodiments will now be described with reference to the figures.

[0033] Fig. 1 shows a schematic diagram of a processing device 10 in which various examples of the present technology may be embodied. The device comprises a data processing circuit 12, which performs data processing operations on data items according to instruction sequences executed by the circuit. These instructions are retrieved from a memory 14 to which the data processing device has access, and for this purpose a fetch circuit 16 is provided, as will be familiar to those skilled in the art. Furthermore, the instructions retrieved by the fetch circuit 16 are passed to an instruction decoder circuit 18 (also called a decode circuit), which generates control signals arranged to control the configuration and operation of the processing circuit 12, as well as various aspects of the set of registers 20 and the load / store unit 22. In general, the data processing circuit 12 may be arranged in a pipelined manner, the details of which are not relevant to the present technology. The general arrangement depicted in Fig. 1 is well known to those skilled in the art, and a further detailed description is dispensed with solely for reasons of brevity. As seen in FIG. 1, the registers 20 each comprise storage for multiple data elements such that a processing circuit can apply a data processing operation to a specified data element in a specified register, or to a group of specified data elements ("vectors") in a specified register. Within the set of available registers 20, some are designated as general purpose registers configured to hold data values ​​of a given size that characterize the data processing characteristics of the device. For example, these may be 64-bit data values, although the present technology is not limited to any particular such data value size. Other registers within the set of available registers may be explicitly configured as vector registers, where the vector lengths they hold may further be "scalable," i.e., not fixed, but may vary between predefined upper and lower bounds. For example, the device 10 may be configured according to the Scalable Vector Extension (SVE) of the Arm AArch64 architecture provided by Arm Limited of Cambridge, UK, which supports vector lengths that can vary from a minimum of 128 bits to a maximum of 2048 bits in 128-bit increments.Other registers in the set may be designated as predicate registers, and in the context of a scalable vector processing configuration, as scalable predicate registers configured to hold predicate indicators to control the application of vector processing operations to corresponding scalable vectors. When configured in the SVE configuration described above, the device 10 includes 32 scalable vector registers and 16 scalable predicate registers (in addition to other general purpose registers). Other examples and types of registers may be provided, but are not directly relevant to the present technology and will not be described herein. The use of these (scalable) vector (predicate) registers with respect to data elements held in registers 20 will be described in more detail below with reference to some specific embodiments. Data values ​​required by data processing circuit 12 in the execution of instructions, and data values ​​generated as a result of those data processing instructions, are written to and read from memory 14 by load / store unit 22. It should also be noted that, in general, memory 14 of FIG. 1 can be viewed as one example of a computer-readable storage medium capable of storing instructions of the present technology, typically as part of a predefined sequence of instructions ("program") that the processing circuitry subsequently executes. However, the processing circuitry may access such programs from a variety of different sources, such as RAM, ROM, a network interface, etc. This disclosure describes various novel instructions that the processing circuitry 12 can execute, and the following figures provide further explanation of the nature of these instructions, variations in data processing circuitry to support execution of these instructions, etc.

[0034] FIG. 2A illustrates the use of vector processing instructions to specify vector processing operations utilizing predicate information encoded using a predicate counting technique in some examples. A vector processing instruction 80 is shown comprising an opcode 81, an indication of a source operand 82, and an indication of a source predicate register 83. A predicate data value 84 retrieved from the source predicate register 83 includes an element count indication 85 and an element size indication 86. Thus, the predicate information represented by the predicate data value 84 represents a sequence of predicate indicators, each of size 86, repeating count 85 times. For example, if size 86 indicates that the predicate indicator corresponds to an element size of 8 bits, and count 85 represents "8", this indicates predicate information covering a 64-bit width (i.e., element size x element count). An example set of predicate indicators 87 is shown, each labeled "1" to indicate an active set of elements. This set of predicate indicators 87 then controls which elements of a data item (retrieved from an indicated source 82 in the instruction 80) undergo the vector processing operation. The vector processing operation is defined by a particular opcode 82 in the instruction 80. The result 90 of the vector processing may then be used in various ways (e.g., stored in a destination register and also specified in the instruction, but that aspect is omitted from the diagram of FIG. 2A merely for clarity). Thus, a vector processing circuit 89 is provided as part of the processing circuit 12 shown in FIG. 1, and the decoding of the predicated data value is also performed within this processing circuit 52. It should also be understood that the example shown in FIG. 2A is an example of a particularly short and simple set of eight identical predicate indicators 87, and this is done merely for clarity of explanation. More generally, the technique may be readily used to encode much longer sequences of identical predicate indicators, such as are applicable to examples of scalable vector processing such as those according to the above-mentioned SVE of the Arm AArch64 architecture.

[0035] 2B illustrates the use of vector processing instructions to specify vector processing operations that utilize predicate information encoded using a predicate counting technique in some examples. A vector processing instruction 250 is shown comprising an opcode 251, an indication of a destination operand 252, an indication of a source operand 253, and an indication of a source predicate register 254. A predicate data value 255 retrieved from the source predicate register 254 includes an element count indication 256 and an element size indication 257. Thus, the predicate information represented by the predicate data value 255 represents a sequence of predicate indicators, each of size 257, repeated count 256 times. For example, if size 257 indicates that the predicate indicator corresponds to an element size of 16 bits and count 256 represents "32", this indicates predicate information covering a width of 512 bits (i.e., element size x element count). The operation defined by a particular opcode 251 in instruction 250 in the case of Figure 2B is either a vector load or vector store operation, and is executed by load / store circuitry 260 (shown separately in Figure 1 for clarity of illustration, but which may generally be considered part of the programmed processing circuitry). In the case of a load operation, a source operand 253 indicates a range of memory locations, and a set of predicate indicators 258 controls which "elements" are loaded from that range of memory locations, which are then loaded into a destination vector register 252. The load / store circuitry 260 accesses memory 262 and registers 261 (e.g., memory 14 and registers 20 in Figure 1). In the case of a store operation, a source operand 253 indicates a source vector register, and a set of predicate indicators 258 controls which elements from that source are stored into the range of memory locations indicated by destination operand 252. For both load and store operations, a range of memory locations may be indicated by a set of pointers (i.e., the source or destination operand associated with a memory location is actually a vector register holding a set of pointers), allowing gather / scatter type loads and stores to be performed (predicated by a predicate data value).

[0036] 3 illustrates generally the use of vector processing instructions that specify vector processing operations that utilize predicate information encoded using a predicate counting technique in some examples. A vector processing instruction 30 is shown that includes an opcode 32, an indication of a first source vector register 34, an indication of a second source vector register 36, and an indication of a source predicate register 38. A predicate data value 40 retrieved from the source predicate register 38 includes an element count indication 42 and an element size indication 44. Thus, the predicate information represented by the predicate data value 40 represents a sequence of predicate indicators, each of size 44, that repeat count 42 times. For example, if size 44 indicates that the predicate indicator corresponds to an 8-bit element size, then count 42 representing "8" indicates predicate information covering a 64-bit width (i.e., element size x element count). An example set of predicate indicators 46 is shown, each labeled "1" to indicate an active set of elements. This set of predicate indicators 46 then controls which elements of the respective vectors 48 and 50 (fetched from the two source vector registers according to the instructions 34 and 36 in the instruction 30) undergo a vector processing operation. This vector processing operation is defined by a particular opcode 32 in the instruction 30. The result of the vector processing may be one or more result vectors 54. In fact, typically one or more destination vector registers are also specified in the instruction, but that aspect is omitted in the diagram of FIG. 3 merely for clarity. Thus, the vector processing circuit 52 is provided as part of the processing circuit 12 shown in FIG. 1, and the decoding of the predicate data values ​​is also performed within this vector processing circuit 52. It should also be understood that the example shown in FIG. 3 is an example of a particularly short and simple set of eight identical predicate indicators 46, and this is done merely for clarity of explanation. More generally, the technique may be readily used to encode much longer sequences of identical predicate indicators, such as are applicable to examples of scalable vector processing such as those according to the above-mentioned SVE of the Arm AArch64 architecture.

[0037] FIG. 4A illustrates the use of inversion indicators in predicate encoding, according to some examples. In addition to the element count 64 and element size 66 described above, the predicate data value 60 further includes an inversion indication 62 (which may be provided as a single bit indicating inversion or non-inversion). The effect of the inversion indication 62 is to invert the relative positions of the active and inactive predicate elements. Thus, in a first configuration 68, an active element is followed by an inactive element, and in a second configuration 70, an inactive element is followed by an active element. Once again, it should be understood that the example shown in FIG. 4A is also a particularly short and simple set of predicate indicators, which is done merely for clarity of explanation. More generally, the technique may easily represent a much longer sequence of identical active and / or inactive elements of a glass. From a slightly different perspective, it can be seen that the value of the inversion indication 62 indicates the start of the active elements, whereby in one configuration, the active elements start at the leftmost predicate position and repeat according to the count value, with the remaining predicate indications being the inactive elements. In another configuration, the inactive element starts at the leftmost predicate position and repeats until it reaches a position that allows the remainder of the active element's predicate instructions to have a multiplicity that matches the count value.

[0038] 4B shows a schematic representation of the variable position of the boundary between the element count indication and the element size indication in the encoding used for the predicate count definition according to various examples. The encoding shown includes an inversion indicator (I), an element count indicator, and an element size indicator. Furthermore, as shown in the figure, the boundary between the part carrying the element count indicator and the part carrying the element size indicator can vary in position. In particular, it varies depending on the element size shown, which allows more or less space for the element count indicator.

[0039] FIG. 4C shows diagrammatically an example of a bit value determining the location of the boundary between an indication of element count and an indication of element size in the encoding used for the predicate count definition according to various examples. Here, the encoding used represents the element size as "1000". The leading 1 therefore determines the indicated value and therefore this "active bit" determines the boundary between the part holding the element count indicator and the part holding the element size indicator. If a different element size were indicated by "10", this would move the boundary to the right (in the orientation shown), allowing two more bits to be used for the element count indicator.

[0040] FIG. 5 illustrates diagrammatically the encoding used for the predicate count definition according to various examples. This particular encoding utilizes the lowest 16 bits of the SVE predicate register. Bit

[15] is used as an inversion indicator, with 1 indicating inversion and 0 indicating non-inversion. Bits [14:0] are used to represent the element count and element size, with the boundary between the element count portion and the element size portion at a variable position depending on the element size indication. As seen in FIG. 4B, the encoding used to represent the element size is a binary 1 followed by 0 to 4 binary 0s. As shown in the figure, [0]=1 indicates a byte element size (leaving bits [14:1] available to encode an element count), [1:0]=10 indicates a halfword element size (leaving bits [14:2] available to encode the element count), [2:0]=100 indicates the word element size (leaving bits [14:3] available to encode the element count), [3:0]=1000 indicates a doubleword element size (leaving bits [14:4] available to encode the element count), [4:0]=10000 indicates a quadword element size (leaving bits [14:5] available to encode the element count).

[0041] Furthermore, the encoding of [4:0]=00000 is assigned the special meaning of all false (inactive) elements. Thus, when

[15] =1 and [4:0]=00000, this indicates all true (active) elements, i.e., this is the canonical form for all active predicates using the predicate counter representation.

[0042] FIG. 6 illustrates diagrammatically the reading and writing of predicate data values ​​that define a predicate count according to various examples. A predicate register is shown, and a read operation on the predicate register takes only the least significant 16 bits, and any bits above this are ignored. The most significant bit of the predicate register is labeled "MAX" in the figure, since the size of the predicate register may vary, especially in the context of a scalable vector processing implementation. For example, an SVE configuration may be implemented with scalable vector registers that are 128-2048 bits in length and can hold 64, 32, 16, or 8-bit elements. Correspondingly, a scalable predicate register that is 1 / 8 the scalable vector length (to allow 8-bit elements to be predicated) may be 16-256 bits in length. Where the minimum predicate register length is 16 bits, the technique is presented in that context, with a predicate counter encoding that occupies a 16-bit space. Thus, as shown in FIG. 6, any bits above this are ignored when reading the predicate data value. Conversely, when a predicate data value is written to a predicate register, only the least significant 16 bits are set according to the predicate counter encoding, and any bits above this are set to zero.

[0043] 7 shows a schematic of the use of a predicate generating instruction according to various examples. A predicate generating instruction 100 comprises an opcode 102 indicating the type of instruction, a destination predicate register indication 104, an element size indication 106, two source general purpose register indications 107 and 108, and an operand (VL) 110 indicating the number of vectors controlled by this predicate. The specific operations that the predicate generating instruction 100 triggers to generate the contents of the predicate from the contents of the two source general purpose registers are of the range type, the following are just a few examples: If the difference between the second signed scalar operand and the first signed scalar operand is positive or zero, then this difference is used to generate a predicate count. If the difference between the second unsigned scalar operand and the first unsigned scalar operand is positive or zero, then this difference is used to generate a predicate count. If the difference between the first signed scalar operand and the second signed scalar operand is positive or zero, then this difference is used to generate a predicate count. If the difference between the first unsigned scalar operand and the second unsigned scalar operand is positive or zero, then this difference is used to generate a predicate count.

[0044] The operand (VL) 110 indicating the number of vectors to be controlled by this predicate in this example is a single bit indicating either two or four vectors to be controlled. This then determines the maximum value that can be stored in the element count of the predicate mask. For example, for a scalable vector length (SVL) of 512 bits, if four vectors are controlled by the predicate, the total (four times) SVL will be 256 bytes. With a minimum element size of byte length, the maximum count value is 256. With reference to the exemplary encoding of FIG. 5, this requires that bits [8:1] are used for the element count, with bit [0] set to 1 to indicate a byte size element. Bits [14:9] are then ignored in this example (e.g., when the predicate is read for use in controlling a 512-bit SVL). Bit

[15] is still read as an inversion indicator. The predicate is thus generated accordingly by a predicate encoding circuit 112 provided as part of the processing circuit 12 shown in FIG. 1. Thus, the operand (VL) 110 indicating the number of vectors to be controlled by this predicate also determined the number of elements to be considered for the all active and last active checks when setting the condition flags.

[0045] 8 illustrates the use of a size indicator in a vector processing instruction that converts a predicate counter to a predicate mask according to various examples. A predicate conversion instruction 120 comprises an opcode 122 indicating a type of instruction, an indication of a first destination predicate register 124, an indication of a second destination predicate register 125, an element size indication 126, and a source predicate register indication 128. A predicate data value 130 (with a predicate counter encoding) is retrieved from the source predicate register 128 and forms the subject of a predicate decode 132 (performed by the processing circuit 12 shown in FIG. 1). The resulting predicate mask is then stored across two destination predicate registers 134 and 136.

[0046] 9A illustrates the use of sub-portion indicators in predicate conversion instructions according to various examples. The predicate conversion instruction 150 comprises an opcode 151 indicating the type of instruction, a destination predicate register indication 152, an element size indication 153, a source predicate register indication 154, and a sub-portion selection indicator 155. A predicate data value 160 (with predicate counter encoding) is retrieved from the source predicate register 154 and forms the subject of a predicate decode 161 (performed by the processing circuitry 12 shown in FIG. 1). This results in an "extended" predicate mask 162, which has a length corresponding to a full (in this example) four vector lengths. The sub-portion selection indicator 155 controls a selection circuit 163 that causes a selected one of the sub-portions to be stored in the destination predicate register 164.

[0047] 9B shows a schematic of an all-true predicate generation instruction 170 according to various examples, with an opcode 171 indicating the type of instruction, a destination predicate register indication 172, an element size indication 173, and an operand (VL) 174 indicating the number of vectors to be controlled by this predicate. This instruction generates a defined number of vectors corresponding to the active elements in the predicate counter encoding.

[0048] 10 shows a schematic representation of the use of a predicate count instruction according to various examples. A predicate count instruction 180 comprises an opcode 181 indicating the type of instruction, a destination general register indication 182, a source predicate register indication 183, an element size indication 184, and an operand (VL) 185 indicating a limit on the number of elements to be counted, corresponding to the number of vectors controlled by this predicate. A predicate data value 186 (with predicate counter encoding) is retrieved from the source predicate register 186 and forms the subject of a predicate decode 187 (performed by the processing circuitry 12 shown in FIG. 1). An element count circuit 188 (also forming part of the processing circuitry 12 shown in FIG. 1) then determines an element count from the decoded predicate and stores this scalar count value in a general register 189. This count takes into account the element size in the number of vectors indicated.

[0049] Various examples of instructions are given above in relation to predicate counters. In general, predicated multi-vector instructions consume a predicate counter as the dominating predicate. These instructions take into account both the number of active elements and the width (size) of the elements in the predicate counter encoding. This allows the width of the operation to be narrower or wider than the element size in the predicate counter encoding.

[0050] FIG. 11 shows a schematic representation of a simulator implementation that may be used. While the above embodiments implement the invention in terms of apparatus and methods for operating specific processing hardware supporting the technique, it is also possible to provide an instruction execution environment according to the embodiments described herein implemented by the use of a computer program. Such computer programs are often referred to as simulators insofar as they provide a software-based implementation of a hardware architecture. Various simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on a host processor 200, which optionally runs a host operating system 210 and supports a simulator program 220. In some arrangements, there may be multiple layers of simulation between the hardware and the instruction execution environment provided, and / or there may be multiple different instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations that run at reasonable speeds, but such an approach may be justified in certain situations, such as when it is desirable to run code native to another processor for compatibility or reuse reasons. For example, a simulator implementation may provide an instruction execution environment that has additional functionality not supported by the host processor hardware, or that is typically associated with a different hardware architecture. An overview of simulation is given in "Some Efficient Architecture Simulation Techniques", Robert Bedichek, Winter 1990 USENIX Conference, pp. 53-63.

[0051] While embodiments have been described above with reference to particular hardware constructs or features, equivalent functionality may be provided in the simulated embodiments by appropriate software constructs or features. For example, particular circuitry may be implemented as computer program logic in the simulated embodiments. Similarly, memory hardware such as registers or caches may be implemented as software data structures in the simulated embodiments. In configurations in which one or more of the hardware elements referenced in the foregoing embodiments reside on host hardware (e.g., host processor 200), some simulated embodiments may utilize the host hardware where appropriate.

[0052] Simulator program 220 may be stored in a computer-readable storage medium (which may be a non-transitory medium) and provides a program interface (an instruction execution environment) to target code 230 (which may include applications, operating systems, and hypervisors) that is the same as the interface of the hardware architecture being modeled by simulator program 220. Thus, program instructions of target code 230, including the above-mentioned instructions for generating and manipulating the above-mentioned predicates encoded as predicate counters, may be executed from within the instruction execution environment using simulator program 220, such that host computer 200, which does not actually have the hardware characteristics of device 10 described above, can emulate those characteristics.

[0053] In summary, an apparatus, method and program for predication of multiple vectors in vector processing is disclosed. Encoding of predicate information including element size and element count is disclosed, where the predicate information includes multiple consecutive identical predicate indicators given by the element count, each predicate indicator corresponding to an element size.

[0054] In this application, the term "configured to..." is used to mean that elements of an apparatus have a configuration that is capable of performing a defined operation. In this context, "configuration" refers to a manner of arrangement or interconnection of hardware or software. For example, an apparatus may have dedicated hardware that provides the defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not imply that the apparatus elements need to be modified in any way to provide the defined operation.

[0055] Although illustrative embodiments have been described in detail herein with reference to the accompanying drawings, it will be understood that the invention is not limited to exact embodiments thereof, and various changes, additions and modifications may be made by those skilled in the art without departing from the scope and spirit of the invention as defined in the appended claims. For example, various combinations of the features of the following dependent claims may be made with the features of the independent claims without departing from the scope of the invention.

Claims

1. An apparatus comprising: A decoding circuit for decoding an instruction; A processing circuit for applying the vector processing operation specified by the instruction to an input data vector; The decoding circuit, in response to a vector processing instruction, Specifies a vector processing operation, One or more source operands, And a source predicate register, and Generates a control signal for causing the processing circuit to execute the vector processing operation on the source operand, the control signal further causing the processing circuit to selectively apply the vector processing operation to elements of the one or more source operands predicated by a predicate indicator decoded from a predicate data value fetched from the source predicate register; The predicate data value has an encoding, An element size, And an element count indicating a plurality of consecutive identical predicate indicators, each predicate indicator corresponding to the element size, the apparatus.

2. The apparatus according to claim 1, wherein the processing circuit is configured to decode the predicate data value to generate the consecutive identical predicate indicators and a further sequence of identical predicate indicators, the consecutive identical predicate indicators and the further sequence of identical predicate indicators including mutually opposite activity indications.

3. The apparatus according to claim 1, wherein the processing circuit generates all the predicate indicators as the consecutive identical predicate indicators in response to the element count having a predetermined value.

4. The apparatus according to claim 1, wherein the encoding of the predicate data value further includes an inversion bit, and the repeated activity indication forming the consecutive identical predicate indicators depends on the inversion bit.

5. The encoding of the predicate data value uses element size encoding to indicate the element size, and the element size encoding includes an indication that the element size is a byte length, an indication that the element size is a half-word length, an indication that the element size is a word length, an indication that the element size is a double-word length, The apparatus according to claim 1, comprising an indication for at least one of the element size being a quad-word length.

6. The encoding is the predicate data value includes a predetermined portion of the predicate data value used to indicate the element size and the element count, The apparatus according to claim 1, wherein a boundary position between a first sub-portion indicating the element size and a second sub-portion indicating the element count in the predetermined portion of the predicate data value depends on the indicated element size.

7. The bit position of the active bit in the first sub-portion indicates the element size, and the bit position of the active bit defines the boundary position. The apparatus according to claim 6.

8. The encoding of the predicate data value is limited to a predetermined number of bits of the predicate data value, and the processing circuit is configured to ignore any further bits held in the source predicate register that exceed the bits forming the predetermined number of bits of the predicate data value when reading the predicate data value from the source predicate register. The apparatus according to claim 1.

9. The encoding of the predicate data value is limited to a predetermined number of bits of the predicate data value, The apparatus according to claim 1, wherein when the processing circuit writes a new predicate data value to the target predicate register, any additional bits that can be held within the target predicate register beyond the bits forming the predetermined number of bits of the predicate data value are configured to be set to a predetermined value.

10. The apparatus according to claim 1, wherein the decoding circuit generates a control signal for causing the processing circuit to generate a predicate data value indicating a corresponding element size and a corresponding element count in response to a predicate generation instruction that specifies a predicate to be generated and the number of vectors to be controlled by the predicate to be generated.

11. The apparatus according to claim 1, wherein the decoding circuit generates a control signal for causing the processing circuit to generate an all-true predicate data value indicating all active elements for the predicate indicator in response to an all-true predicate generation instruction that specifies all true predicates to be generated.

12. The apparatus according to claim 1, wherein the decoding circuit generates a control signal for causing the processing circuit to generate an all-false predicate data value indicating all non-active elements for the predicate indicator in response to an all-false predicate generation instruction that specifies all false predicates to be generated.

13. The decoding circuit generates a control signal for causing the processing circuit to decode the predicate data value to be converted and generate a converted predicate data value in response to a predicate conversion instruction that specifies a source predicate register holding the predicate data value to be converted, The apparatus according to claim 1, wherein the converted predicate data value includes a direct mask style representation in which the bit value at a bit position indicates the predicate of an element within a target data item.

14. The apparatus according to claim 13, wherein the predicate conversion instruction specifies two or more destination predicate registers, the control signal causes the processing circuit to generate two or more converted predicate data values, each of the two or more converted predicate data values includes the direct mask style representation, and each of the two or more converted predicate data values corresponds to a different subset of the predicate indicators represented by the predicate data values to be converted.

15. The apparatus according to claim 14, wherein the predicate conversion instruction specifies the multiplicity of two or more of the converted predicate data values to be generated.

16. The apparatus according to claim 13, wherein the predicate conversion instruction specifies which of a plurality of possible subsets of the predicate bits represented by the predicate data values to be converted are to be generated.

17. The apparatus according to claim 13, wherein the decoding circuit generates a control signal in response to a predicate count instruction that specifies a source predicate register that holds the predicate data values to be counted, the control signal causes the processing circuit to decode the predicate data values to be converted and determine the predicate indicators indicated by the predicate data values to be converted, and store a scalar value corresponding to the number of active elements in the predicate indicator in a destination general-purpose register.

18. The predicate count instruction specifies an upper limit on the number of active elements to be counted, and the upper limit is one of two vector lengths, and one of four vector lengths, corresponding to the apparatus according to claim 17.

19. The one or more source operands are one source vector register, two source vector registers, or three source vector registers, for the apparatus according to claim 1.

20. The one or more source operands are The apparatus according to claim 1, comprising a range of memory locations.

21. The apparatus according to claim 20, wherein the range of memory locations is indicated by a set of pointers.

22. The apparatus according to claim 21, wherein the vector processing instruction further specifies a destination vector register.

23. The apparatus according to claim 1, wherein the vector processing instruction further specifies a destination memory location.

24. A data processing method, comprising: decoding an instruction; controlling a processing circuit to apply a vector processing operation specified by the instruction to an input data vector, wherein the decoding is responsive to a vector processing instruction to specify a vector processing operation, one or more source operands, and a source predicate register, and generate a control signal for causing the processing circuit to perform the vector processing operation on the one or more source operands, the control signal further causing the processing circuit to selectively apply the vector processing operation to elements of the one or more source operands predicated by a predicate indicator decoded from a predicate data value retrieved from the source predicate register; wherein the predicate data value has an encoding, an element size, and an element count indicating a plurality of consecutive identical predicate indicators, each predicate indicator corresponding to the element size,

25. A computer program for controlling a host processing device to provide an instruction execution environment, comprising: decoding logic for decoding an instruction; processing logic for applying the vector processing operation specified by the command to the input data vector; wherein the decoding logic, in response to a vector processing instruction, specifies a vector processing operation, one or more source operands, and a source predicate register, and generates a control signal for causing the processing logic to execute the vector processing operation with respect to the source operands, the control signal further causing the processing circuit to selectively apply the vector processing operation to elements of the one or more source operands that are predicated by a predicate indicator decoded from a predicate data value fetched from the source predicate register; wherein the predicate data value has an encoding, an element size, and an element count indicating a plurality of consecutive identical predicate indicators, each predicate indicator corresponding to the element size; a computer program. ​