Techniques for performing memory access operations - Patents.com

JP2025504087A5Pending Publication Date: 2025-12-23ARM LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024545864
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-02-07
Filing Date
2022-12-20
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize the address indication capability of vector instructions in a vector processing system, resulting in the inability to efficiently perform vector aggregation or decentralization operations, and security is difficult to guarantee.

Method used

By using vector instructions containing address indication and constraint information, capabilities in multiple vector registers are determined based on these capabilities to determine the storage address of each data element and security checks are performed to perform vector aggregation or decentralization operations.

Benefits of technology

It realizes efficient execution of vector aggregation or decentralized operations in vector processing systems, while improving the security of the system and the flexibility of address indication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An apparatus is described having a processing circuit for performing vector processing operations, a set of vector registers, and an instruction decoder for decoding the vector instructions to control the processing circuit to perform the required operations. The instruction decoder is responsive to a given vector memory access instruction specifying a plurality of memory access operations, each memory access operation being executed to access an associated data element to determine from a data vector designation field of the given vector memory access instruction at least one vector register in the set of vector registers associated with the plurality of data elements, and to determine from at least one capability vector designation field of the given vector memory access instruction a plurality of vector registers in the set of vector registers including a plurality of capabilities. Each capability is associated with one of the data elements in the plurality of data elements and provides an address designation and constraint information that constrains the use of the address designation when accessing the memory. The number of vector registers determined from the at least one capability vector designation field is greater than the number of vector registers determined from the data vector designation field. The instruction decoder controls the processing circuitry to determine, for each given data element within the plurality of data elements, a memory address based on the addressing instructions provided by the associated capability, determine whether a memory access operation used to access the given data element is permitted for the determined memory address taking into account the constraint information of the associated capability, and enable execution of the memory access operation for each data element for which the memory access operation is permitted.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present techniques relate to the field of data processing, and more particularly, to manipulating memory access operations.

[0002] Vector processing systems have been developed that attempt to improve code density, and often performance, by allowing a given vector instruction to be executed to perform an operation defined by the given vector instruction independently on multiple data elements in a vector of data elements. Thus, in the context of memory access operations, it is possible to load multiple consecutive data elements from memory into a specified vector register in response to a vector load instruction, or to store multiple consecutive data elements from a specified vector register into memory in response to a vector store instruction. It is also possible to provide vector aggregation or vector scatter variants of those vector load or vector store instructions to allow the data elements to be processed to reside anywhere in memory. When using such vector aggregation or vector scatter instructions, in addition to a vector being identified for the multiple data elements to be processed, the vector can also be identified to provide multiple address indications used to determine the memory address of each data element.

[0003] There is growing interest in capability-based architectures, where certain capabilities are defined for a given process and an error can be triggered if there is an attempt to perform an operation outside the defined capabilities. Capabilities can take a variety of forms, but one type of capability is a bounded pointer (which may also be referred to as a "fat pointer").

[0004] Each capability may include constraint information used to limit the operations that may be performed when using that capability. For example, considering a bounded pointer, this may provide information used to identify a non-extensible range of memory addresses accessible by the processing circuit when using that capability, along with one or more permission flags that identify associated permissions.

[0005] To benefit from the security benefits provided through the use of capabilities, it is desirable to support the execution of vector aggregate or vector scatter instructions, but allow various address directives to be specified by the capabilities. However, capabilities that provide address directives are inherently larger than equivalent standard address directives due to constraint information that is provided in association with the address directives to form the capabilities. Summary of the Invention

[0006] In a first exemplary configuration, an apparatus is provided, the apparatus comprising: a processing circuit for performing vector processing operations; a set of vector registers; and an instruction decoder for decoding the vector instruction and controlling the processing circuit to perform the vector processing operations specified by the vector instruction, the instruction decoder being responsive to a given vector memory access instruction specifying a plurality of memory access operations, each memory access operation being performed to access an associated data element, to determine from a data vector indication field of the given vector memory access instruction at least one vector register in the set of vector registers associated with the plurality of data elements, and to determine a plurality of vector registers in the set of vector registers including a plurality of capabilities, each capability being associated with one of the data elements in the plurality of data elements, and to provide an address indication and a capability constraining use of the address indication when accessing the memory. and constraint information, wherein the number of vector registers determined from the at least one capability vector indication field is greater than the number of vector registers determined from the data vector indication field, and the instruction decoder is further configured to control the processing circuitry to: determine, for each given data element within the plurality of data elements, a memory address based on the address indication provided by the associated capability; determine whether a memory access operation used to access the given data element is permitted for the determined memory address taking into account the constraint information of the associated capability; enable performance of the memory access operation for each data element for which the memory access operation is permitted, and wherein performance of the memory access operation for any given data element moves the given data element between the determined memory address in memory and the at least one vector register.

[0007] In a further exemplary arrangement, there is provided a method of performing memory access operations in an apparatus providing processing circuitry for performing vector processing operations and a set of vector registers, the method comprising: in response to a given vector memory access instruction specifying a plurality of memory access operations, each memory access operation being performed to access an associated data element, determining from a data vector indication field of the given vector memory access instruction at least one vector register in the set of vector registers associated with the plurality of data elements, determining a plurality of vector registers in the set of vector registers comprising a plurality of capabilities, each capability being associated with one of the data elements in the plurality of data elements, providing address indications and constraint information constraining use of the address indications when accessing the memory, and determining a plurality of vector registers in the set of vector registers comprising a plurality of capabilities, each capability being associated with one of the data elements in the plurality of data elements, and employing an instruction decoder, wherein a number of vector registers determined from the capability vector designation field is greater than a number of vector registers determined from the data vector designation field; and causing the processing circuitry to determine, for each given data element within the plurality of data elements, a memory address based on the address designation provided by the associated capability, determine whether a memory access operation used to access the given data element is permitted for the determined memory address taking into account constraint information of the associated capability, enable performance of the memory access operation for each data element for which the memory access operation is permitted, and wherein performance of the memory access operation for any given data element moves the given data element between the determined memory address in memory and the at least one vector register.

[0008] In another exemplary configuration, a computer program for controlling a host data processing apparatus to provide an instruction execution environment, the computer program including: processor program logic for performing vector processing operations; vector register emulation program logic for emulating a set of vector registers; and instruction decode program logic for decoding the vector instruction and controlling the processor program logic to perform the vector processing operations specified by the vector instruction, the instruction decode program logic is responsive to a given vector memory access instruction specifying a plurality of memory access operations, each memory access operation being performed to access an associated data element, to determine from a data vector designation field of the given vector memory access instruction at least one vector register in the set of vector registers associated with the plurality of data elements, and to determine a plurality of vector registers in the set of vector registers including a plurality of capabilities, each capability being associated with a data element in the plurality of data elements. the capability vector indication field being associated with one of the capability vector indication fields and providing an address indication and constraint information constraining the use of the address indication when accessing the memory, the number of vector registers determined from the at least one capability vector indication field being greater than the number of vector registers determined from the data vector indication field, and the instruction decode program logic is further configured to control the processing program logic to: determine, for each given data element within the plurality of data elements, a memory address based on the address indication provided by the associated capability; determine whether a memory access operation used to access the given data element is permitted with respect to the determined memory address taking into account the constraint information of the associated capability; enable execution of the memory access operation for each data element for which the memory access operation is permitted; and for any given data element, execution of the memory access operation causes the given data element to move between the determined memory address in the memory and the at least one vector register.

[0009] In yet another exemplary configuration, an apparatus is provided comprising: processing means for performing vector processing operations; a set of vector register means; and instruction decode means for decoding the vector instruction and controlling the processing means to perform the vector processing operation specified by the vector instruction, wherein the instruction decoder means, in response to a given vector memory access instruction specifying a plurality of memory access operations, each memory access operation being performed to access an associated data element, determines from a data vector indication field of the given vector memory access instruction at least one vector register means in the set of vector register means associated with the plurality of data elements, determines a plurality of vector register means in the set of vector register means comprising a plurality of capabilities, each capability being associated with one of the data elements in the plurality of data elements, and determines an address indication when accessing the memory from the data vector indication field of the given vector memory access instruction. and constraint information constraining use of the capability vector means, wherein the number of vector register means determined from the at least one capability vector indication field is greater than the number of vector register means determined from the data vector indication field, and the instruction decode means is further configured to control the processing means to: determine, for each given data element within the plurality of data elements, a memory address based on the address indication provided by the associated capability, determine whether a memory access operation used to access the given data element is permitted for the determined memory address taking into account the constraint information of the associated capability, enable performance of the memory access operation for each data element for which the memory access operation is permitted, and wherein performance of the memory access operation for any given data element moves the given data element between the determined memory address in memory and the at least one vector register means. [Brief description of the drawings]

[0010] The present technique will now be further described, by way of example only, with reference to examples of the technique illustrated in the accompanying drawings, in which: [Figure 1]FIG. 1 is a block diagram of an apparatus according to an exemplary implementation. [Diagram 2] 1 illustrates the use of tag bits associated with capabilities, according to one exemplary implementation. [Figure 3A] A diagram illustrating different ways in which a valid capability indication (in one example, taking the form of a tag bit) can be stored in association with each capability-sized block of a vector register to indicate whether the capability-sized block stores a valid capability, according to an exemplary implementation. [Figure 3B] A diagram illustrating different ways in which a valid capability indication (in one example, taking the form of a tag bit) can be stored in association with each capability-sized block of a vector register to indicate whether the capability-sized block stores a valid capability, according to an exemplary implementation. [Figure 4A] 1 is a flow diagram illustrating how tag bits maintained in association with each capability-sized block of a vector register may be managed, according to an example implementation. [Figure 4B] 1 is a flow diagram illustrating how tag bits maintained in association with each capability-sized block of a vector register may be managed, according to an example implementation. [Figure 5A] FIG. 1 illustrates fields that may be provided in a vector memory access instruction, according to an exemplary implementation. [Figure 5B] 1 is a flow diagram illustrating steps performed when executing such a vector memory access instruction, according to one exemplary implementation. [Figure 6A] 1 is a flow diagram illustrating a technique that may be used to determine a number of vector registers that hold necessary capabilities to be used when performing aggregation and distribution operations, according to an example implementation. [Figure 6B]1 is a flow diagram illustrating a technique that may be used to determine a number of vector registers that hold necessary capabilities to be used when performing aggregation and distribution operations, according to an example implementation. [Figure 7] 1 illustrates a schematic of how a set of vector registers may be logically partitioned into multiple sections, according to one exemplary implementation. [Figure 8A] FIG. 2 illustrates a particular exemplary configuration of data elements and associated capabilities that may be used when performing aggregation or distribution operations of the type described herein. [Figure 8B] FIG. 2 illustrates a particular exemplary configuration of data elements and associated capabilities that may be used when performing aggregation or distribution operations of the type described herein. [Figure 8C] FIG. 2 illustrates a particular exemplary configuration of data elements and associated capabilities that may be used when performing aggregation or distribution operations of the type described herein. [Figure 9] 4 is a flow diagram illustrating how associated capabilities may be determined for each data element according to one example implementation. [Figure 10] FIG. 13 illustrates an example of overlapped execution of vector instructions. [Figure 11] We present three examples of scaling the amount of overlap between consecutive vector instructions between different processor implementations, or between different instances of execution of the instructions at run-time. [Figure 12] FIG. 10 is a flow chart showing how a sequence of vector capability memory transfer instructions may be used in one exemplary implementation to move capabilities between memory and vector registers to ensure that the capabilities are stored in a configuration within vector registers that enables their use when performing aggregation and distribution operations in the manner described herein. [Figure 13]1 illustrates generally how different memory banks may be accessed when employing a sequence of vector capability memory transfer instructions to transfer capabilities between memory and vector registers in accordance with the techniques described herein. [Figure 14] 1 shows an example simulator that can be used. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] In accordance with the techniques described herein, an apparatus is provided having a processing circuit for performing vector processing operations, a set of vector registers, and an instruction decoder for decoding the vector instruction to control the processing circuit to perform the vector processing operations specified by the vector instruction. The vector processing operations specified by the vector instruction may be implemented by independently performing the required operations on each of multiple data elements in the vector, which may be performed in parallel, sequentially one after the other, or in groups (e.g., operations within a group may be performed in parallel and each group may be performed sequentially).

[0012] The instruction decoder may be configured to process a given vector memory access instruction that specifies multiple memory access operations, each of which is to be performed to access an associated data element, and thus the multiple memory access operations may be collectively viewed as implementing the vector memory access operation specified by the vector memory access instruction. In particular, in response to such a given vector memory access instruction, the instruction decoder may be configured to determine, from a data vector designation field of the given vector memory access instruction, at least one vector register in a set of vector registers associated with the multiple data elements. Thus, each vector register determined from the data vector designation field may, for example, form a source register for a vector distribution operation that seeks to store a data element from its source register to various locations in memory, or may serve as a destination register for a vector aggregation operation that seeks to load data elements from various locations in memory for storage in its vector register.

[0013] The instruction decoder is also configured to determine, from at least one capability vector designation field of a given vector memory access instruction, a plurality of vector registers in the set of vector registers that include a plurality of capabilities. In one exemplary implementation, a single capability vector designation field is used, and the plurality of vector registers are determined from information in the single capability vector designation field. However, in alternative implementations, multiple capability vector designation fields may be provided, for example, to enable each capability vector designation field to identify a corresponding vector register. In one exemplary implementation, each vector register of the plurality of vector registers includes a plurality of capabilities, and in another example, each vector register of the plurality of vector registers includes a single capability.

[0014] Each capability in the determined plurality of vector registers is associated with one of the data elements in the plurality of data elements and provides an address instruction and constraint information that constrains the use of the address instruction when accessing memory. The constraint information may take various forms, but may, for example, identify range information used to determine an allowable range of memory addresses that may be accessed when using the address instruction provided by the capability, and / or one or more permission attributes that specify the type of access that may be performed using the address instruction (e.g., whether read access is permitted, whether write access is permitted, whether the capability can be used to generate a memory address for an instruction to be fetched and executed, whether access is permitted from a particular level of security or privilege, etc.). In a further example, the constraint information may be a constraint identification value that indicates an entry in a set of constraint information. Each entry in the set of constraint information may take a variety of forms, but may, for example, identify range information used to determine an allowable range of memory addresses that may be accessed when using the address instruction provided by the capability, and / or one or more permission attributes that specify the type of access that may be performed using the address instruction (e.g., whether read access is allowed, whether write access is allowed, whether the capability can be used to generate memory addresses for instructions to be fetched and executed, whether access is allowed from a particular level of security or privilege, etc.). In some implementations, the generated memory address may be a physical memory address that directly corresponds to a location in the memory system, while in other implementations, the generated memory address may be a virtual address for which address translation may need to be performed to determine the physical memory address to be accessed.

[0015] In accordance with the techniques described herein, the number of vector registers determined from the at least one capability vector indication field is greater than the number of vector registers determined from the data vector indication field.

[0016] The instruction decoder is further configured to control the processing circuit to determine, for each given data element in the plurality of data elements, a memory address (which may be either a virtual address or a physical address) based on the address indication provided by the associated capability, and to determine whether a memory access operation used to access the given data element is permitted with respect to that determined memory address, taking into account the constraint information of the associated capability. As mentioned above, the constraint information may take a variety of forms, and thus the checks performed here to determine whether a memory access operation used to access the given data element is permitted may take a variety of forms. Thus, these checks may, for example, identify whether the determined memory address can be accessed given any range constraint information in the capability, but may also determine whether the type of access is permitted (e.g., if the access operation is to perform a write to memory, then the constraint information in the capability permits such a write to be performed).

[0017] The processing circuitry may then be configured to enable execution of a memory access operation for each data element for which the memory access operation is permitted, with execution of the memory access operation for any given data element moving the given data element between the determined memory address in memory and at least one vector register (the direction of the move being understood to depend on whether data is being loaded from memory to a register or stored from a register to memory). In one exemplary implementation, the given data element in its original location may be left untouched during this process, and thus, in that case, the move operation may be performed by copying the given data element. This may, for example, typically be the case at least when loading a data element from memory for storage in a vector register, in which case the data element stored in the vector register is a copy of the data element stored in memory.

[0018] In one exemplary implementation, memory access operations may be performed for each data element for which those memory access operations are permitted, while in other implementations, it may be decided to suppress execution of one or more permitted memory access operations if another of the memory access operations is not permitted. Which permissible access is suppressed in such a situation may depend on the implementation and may depend on where in the vector of data elements is a data element with an associated access that is not permitted. As a purely illustrative example, various accesses may be performed sequentially, so that if one access that is not permitted is detected, it may be decided to suppress the subsequent access, with the previous access already being performed, whether permitted or not.

[0019] In one exemplary implementation, a mechanism is provided for tracking valid capabilities stored in vector registers. In particular, in one exemplary implementation, the apparatus further comprises a capability indication storage providing a valid capability indication field associated with each capability size block in a given vector register of the set of vector registers, each valid capability indication field configured to be set to indicate when the associated capability size block stores a valid capability and is otherwise cleared. In one exemplary implementation, any of the vector registers in the set of vector registers may be capable of storing capabilities, while in another exemplary implementation, the ability to store capabilities may be limited to a subset of the vector registers in the set, in which case the capability indication storage need only provide a valid capability indication field for each capability size block in that subset of vector registers.

[0020] In one exemplary implementation, the capability indication storage may be provided separate from the set of vector registers, while in alternative exemplary implementations the capability indication storage may be incorporated within the set of vector registers.

[0021] To constrain how the valid capability indication field is set, the processing circuitry may be configured to only allow any valid capability indication field to be set to indicate that a valid capability is stored in the associated capability-sized block in response to execution of one or more specific instructions in the set of instructions executable by the device. By restricting the setting of the valid capability indication field in this manner, security may be improved, for example, by prohibiting attempts to indicate that a capability-sized block of generic data in a vector register should be treated as a capability. Thus, operations performed on a vector that do not create a valid capability, either through a non-capability operation or through changing a capability so that it is no longer valid, may be configured to cause the associated valid capability indication field to be cleared, thus indicating that no valid capability is stored therein. Thus, by way of example, a partial write or a non-capability write to a capability-sized data block clears the associated valid capability indication field. The capability indication field may also be cleared by various non-instruction operations, for example, stacking and clearing vector register states associated with exception handling, or in some implementations, a reset operation.

[0022] As mentioned above, the number of vector registers used to provide the capabilities required when executing a given vector memory access instruction as described above is greater than the number of vector registers that contain the data elements that are subject to the memory access operation. In one exemplary implementation, the number of vector registers forming the plurality of vector registers, as determined from at least one capability vector indication field, is a power of two. In particular, the number of vector registers required to store a capability depends on the difference in size between the data elements and the capability, which in one exemplary implementation can vary by a power of two. It is noted herein that when considering the size of a capability, any associated flags (such as the valid capability indication field described above) used to indicate that the capability is a valid capability are not considered part of the capability itself.

[0023] As mentioned above, if desired, multiple capability vector designation fields can be used to designate various vector registers that store capabilities required when executing a given vector memory access instruction. Such an approach allows various vector registers to be arbitrarily positioned relative to one another and designated in the instruction encoding. However, in one exemplary implementation, at least one capability vector designation field is a single capability vector designation field configured to identify one vector register, and the instruction decoder is configured to determine the remaining vector registers of the multiple vector registers based on the determined relationship. Such an approach can be advantageous from the standpoint of instruction encoding, since instruction encoding space is typically very limited and it may not be practical to provide multiple capability vector designation fields to identify each of the vector registers that will store the required capabilities.

[0024] The manner in which the remaining vector registers are determined based on the identified one vector register and the determined relationship can take a variety of forms depending on the implementation. For example, the determined relationship may specify that the vector registers are contiguous with one another, that the vector registers are even / odd pairs, or that there is a known offset between the various vector registers. Alternatively, any other suitable indicated relationship may be used.

[0025] In one particular exemplary implementation, the number of vector registers in the plurality of vector registers that store the required capabilities is 2 N and the single capability vector designation field indicates a first vector register number identifying one vector register, the first vector register number being constrained to have its N least significant bits at a logic 0 value. The instruction decoder is then configured to generate vector register numbers for each of the remaining vector registers by reusing the first vector register number and selectively setting at least one of the N least significant bits to a logic 1 value. This may provide a particularly simple and efficient mechanism for computing the various vector registers that provide the capabilities required when executing a given vector memory access instruction.

[0026] In some implementations, the number of vector registers required to hold a capability is fixed, for example because a given vector memory access instruction is only supported for use with data elements of a particular fixed size, and the capabilities are also fixed-sized, but in the more general case, the number of vector registers can be inferred at run-time by the instruction decoder based on knowledge of the size of the data elements and the size of the capabilities on which a given vector memory access instruction is to be performed.

[0027] There are several ways in which the single capability vector designation field may be configured to indicate the first vector register number. The single capability vector designation field may directly identify the first vector register number in one exemplary implementation, while in other implementations it may specify enough information to allow that first vector register number to be determined. For example, in the above case where the first vector register number is constrained to have its N least significant bits at a logic 0 value, those least significant N bits need not be identified in the single capability vector designation field, but instead may be hardwired to a logic 0 value.

[0028] The manner in which capabilities associated with various data elements are arranged within the vector registers used to provide the capabilities may vary depending on the implementation. However, in one exemplary implementation, for any given pair of data elements associated with adjacent locations in at least one vector register, the associated capabilities are stored in different ones of the plurality of vector registers. It has been found that such an arrangement can enable efficient implementation when executing a given vector memory access instruction.

[0029] The manner in which the location of the associated capability within the plurality of vector registers for any particular data element is determined may vary depending on the implementation. However, in one exemplary implementation, the at least one vector register determined from the data vector designation field comprises a single vector register, and each data element is associated with a corresponding data lane of the single vector register. Furthermore, each capability is located within a capability lane within one vector register of the plurality of vector registers. Note that the width of the data lane is typically different from the width of the capability lane due to the fact that the data elements and capabilities are of different sizes. In such a configuration, for a given data element, the vector register within the plurality of vector registers that contains the associated capability may be determined according to a given number of least significant bits of the lane number of the corresponding data lane, and the capability lane that contains the associated capability may be determined according to the remaining bits of the lane number of the corresponding data lane. This thus provides a particularly efficient mechanism for determining the location of the associated capability for each data element.

[0030] In one particular exemplary configuration, the number of vector registers containing a plurality of capabilities is P, logically considered as a sequence having values ​​0 to P-1, and the number of capability lanes in any given vector register is M, having values ​​0 to M-1. Furthermore, the data lane associated with a given data element is data lane X, having values ​​from 0 to X-1. Using such terminology, in one exemplary implementation, the location of an associated capability within the plurality of vector registers may be determined by dividing X by P to generate a quotient and a remainder, where the quotient identifies the capability lane that contains the associated capability, and the remainder identifies the vector register within the plurality of vector registers that contains the associated capability. Thus, in such an implementation, for a given data element, both the vector register and the capability lane needed to locate the associated capability can be easily and efficiently determined.

[0031] It should be noted that although in the above example the vector registers containing the capabilities are logically considered as a sequence having values ​​0 to P-1, this does not imply that the logical vector numbers associated with these vector registers need to be consecutive logical vector numbers, nor in fact does it imply that the vector registers must be physically located consecutively with respect to each other within the set of vector registers.

[0032] In one exemplary implementation, the set of vector registers may be logically partitioned into multiple sections, each section including a corresponding portion from each of the vector registers in the set of vector registers, and the multiple capabilities may be arranged in the multiple vector registers such that for each data element, the associated capabilities are stored in the same section as the data element. Such an approach may allow the execution of a given vector memory access instruction to be split into multiple "beats", during which only one section of the set of vector registers is accessed to execute the given vector memory access instruction. By allowing the vector memory access instruction to be split into multiple beats, the execution of the vector memory access instruction may be allowed to be overlapped with the execution of one or more other instructions, resulting in a very efficient implementation. In particular, during any particular beat, all of the data elements and capabilities required to execute the memory access operation during that beat may be obtained from a single section of the set of vector registers, leaving any other sections available for access during execution of the overlapped instruction.

[0033] In one example implementation, the processing circuitry may be configured to perform memory access operations over one or more beats for data elements in a given section before performing memory access operations over one or more beats for data elements in the next section. In one example implementation, each beat in the multiple beats used to execute a given vector memory access instruction may access a different section, although this is not a requirement and in some implementations two or more of the beats may access the same section.

[0034] There are several ways in which the capabilities required when executing a given vector memory access instruction as described above may be loaded from memory and then configured within the vector registers in the configuration described above, and indeed, there are several ways in which those capabilities within the vector registers may be stored back to memory at a given time. However, in one exemplary implementation, the instruction decoder is configured to decode a vector capability memory transfer instruction that causes the instruction decoder to control the processing circuitry to transfer the capabilities between the memory and the vector registers, and to reconfigure the capabilities during the transfer such that the capabilities are stored sequentially in the vector registers and the capabilities are deinterleaved across the vector registers, such that any given pair of capabilities within the capabilities stored sequentially in the memory are stored in different ones of the vector registers.

[0035] It should be noted that the vector capability memory transfer instructions used to perform the above steps do not have to directly follow each other and therefore do not have to be executed one after the other. Instead, there may be several separate instructions each performing part of the necessary work, and when all of the instructions are executed, the relocation of capabilities required as they are moved (copied in one example) between memory and the vector registers is performed. The vector capability memory transfer instructions may be either load instructions used to load capabilities from memory into the vector registers, or store instructions used to store capabilities from the vector registers back to memory.

[0036] In one exemplary implementation, each vector capability memory transfer instruction is configured to identify a different capability relative to each other vector capability memory transfer instruction, and each vector capability memory transfer instruction is configured to identify an access pattern that causes the processing circuitry to transfer the identified capability while performing the reconfiguration specified by the access pattern. Thus, in such a configuration, execution of each individual vector capability memory transfer instruction performs the necessary reconfiguration with respect to the capability transferred by that instruction, and then other vector capability memory transfer instructions are used to transfer other capabilities and perform the necessary reconfiguration for those capabilities.

[0037] In such an implementation, various different instructions can be configured to all transfer the same maximum amount of data, which is selected with consideration of the finite memory bandwidth available in any particular system. Such an approach can avoid any individual instruction from stalling, and therefore no ordering state machine is required to implement such an approach. Such an approach also allows other instructions to be scheduled while this capability transfer process is ongoing. Furthermore, by configuring each of the instructions to operate with different capabilities in the manner described above, any individual instruction can be configured for each beat to operate only within the same section of the vector register. As previously described, operating only within a given section allows for overlapping of instructions operating on different sections.

[0038] In one exemplary implementation, the memory is formed from multiple memory banks, and for each vector capability memory transfer instruction, an access pattern is defined such that two or more of the memory banks are accessed when the vector capability memory transfer instruction is executed by the processing circuit. Banked memories make it easier for hardware to implement parallel transfers to / from the memory, and therefore it is beneficial to specify an access pattern that enables this.

[0039] In addition to the vector capability memory transfer instructions described above, vector load and vector store instructions can be used to load data elements from memory into vector registers, or store those data elements from vector registers back to memory, as needed and when required.

[0040] While the number of vector registers used to hold data elements and the number of vector registers used to hold associated capabilities may vary depending on the implementation, in one particular exemplary implementation, the at least one vector register determined from the data vector designation field of a given vector memory access instruction comprises a single vector register, the capability is twice the size of the data elements (as previously mentioned, any flags used to indicate that a capability is a valid capability are not considered to be part of the capability when considering the size of the capability), and the multiple vector registers determined from the at least one capability vector designation field comprise two vector registers. Such an arrangement has been found to provide a particularly useful implementation for performing vector aggregation and distribution operations using memory addresses derived from the capabilities.

[0041] In one exemplary implementation, the given vector memory access instruction may further include an immediate value indicating an address offset, and the processing circuitry may be configured to determine, for each given data element in the plurality of data elements, a memory address of the given data element by combining the address offset with an address indication provided by an associated capability, thereby providing an efficient implementation for calculating memory addresses from address indications provided in various capabilities.

[0042] In one exemplary implementation, a given vector memory access instruction may further include an immediate value indicating an address offset, and for each given data element, the processing circuitry may be configured to update the address instructions of associated capabilities in the multiple vector registers by adjusting the address instructions according to the address offset. Thus, by way of example, once an address instruction in a particular capability is used during execution of a first vector memory access instruction, the address instruction indicated in the capability stored in the vector register may be updated in the manner described above so as to be ready for use in connection with a subsequent vector memory access instruction.

[0043] In some cases, both of the above reconciliation processes may be performed such that an address offset is combined (e.g., added) to the address instruction provided by the capability to identify the memory address to access, and that same updated address is written back to the capability register as an updated address instruction. Typically, the same immediate value is used for both reconciliation processes, although different immediate values ​​may be used for each reconciliation process if desired.

[0044] Specific example implementations will now be described with reference to the accompanying drawings.

[0045] FIG. 1 illustrates generally an example of a data processing apparatus 2 that supports the processing of vector instructions. It will be appreciated that this is a simplified diagram for ease of explanation and that in practice the apparatus may have many elements not shown in FIG. 1 for brevity. The apparatus 2 comprises processing circuitry 4 for performing data processing in response to instructions decoded by an instruction decoder 6. Program instructions are fetched from a memory system 8 and decoded by the instruction decoder to generate control signals that control the processing circuitry 4 to process the instructions in a manner defined by the architecture. For example, the decoder 6 may interpret the opcode of the decoded instruction and any additional control fields of the instruction to generate control signals that cause the processing circuitry 4 to activate appropriate hardware units to perform an operation such as an arithmetic operation, a load / store operation, or a logical operation. The apparatus has a set of scalar registers 10 and a set of vector registers 12. The apparatus may also have other registers (not shown) for storing control information used to configure the operation of the processing circuitry, for example. In response to an arithmetic or logic instruction, the processing circuitry typically reads source operands from registers 10, 12 and writes the result of the instruction back to registers 10, 12. In response to a load / store instruction, data values ​​are transferred between registers 10, 12 and memory system 8 via a load / store unit 18 within processing circuitry 4. Memory system 8 may include one or more levels of data caches and main memory.

[0046] The set of scalar registers 10 includes a number of scalar registers for storing scalar values ​​that include a single data element. Some instructions supported by the instruction decoder 6 and processing circuitry 4 may be scalar instructions that process scalar operands read from the scalar registers 10 to produce scalar results that are written back to the scalar registers.

[0047] The set of vector registers 12 includes several vector registers, each of which is configured to store a vector value including multiple elements. In response to a vector instruction, the instruction decoder 6 may control the processing circuitry 4 to perform several lanes of vector processing on respective elements of a vector operand read from one of the vector registers 12 to generate either a scalar result to be written to the scalar register 10 or a further vector result to be written to the vector register 12. Some vector instructions may generate a vector result from one or more scalar operands, or may perform additional scalar operations on scalar operands in the scalar register file, and may perform lanes of vector processing on vector operands read from the vector register file 12. Thus, some instructions may be mixed scalar vector instructions, where at least one of the one or more source and destination registers of the instruction is the vector register 12 and another of the one or more source and destination registers is the scalar register 10.

[0048] The vector instructions may also include vector load / store instructions that cause data values ​​to be transferred between vector registers 12 and locations in memory system 8. The load / store instructions may include contiguous load / store instructions, where the locations in memory correspond to a contiguous range of addresses, or vector load / store instructions of the aggregate / scatter type, which specify several discrete addresses and control processing circuitry 4 to load data from each of those addresses into respective elements of a vector register, or to store data from respective elements of a vector register into the discrete addresses.

[0049] The processing circuitry 4 may support the processing of vectors having a range of different data element sizes. For example, a 128-bit vector register 12 may be partitioned into sixteen 8-bit data elements, eight 16-bit data elements, four 32-bit data elements, or two 64-bit data elements. A control register may be used to specify the current data element size being used, or alternatively may be a parameter of a given vector instruction being executed.

[0050] The processing circuitry 4 may include several separate hardware blocks for processing different classes of instructions. For example, load / store instructions interacting with the memory system 8 may be processed by a dedicated load / store unit 18, while arithmetic or logical instructions may be processed by an arithmetic logic unit (ALU). The ALU itself may be further divided into a multiply-accumulate unit (MAC) for performing operations including multiplication, and further units for processing other kinds of ALU operations. A floating point unit may also be provided for processing floating point instructions. Pure scalar instructions that do not involve vector processing may be processed by a separate hardware block compared to vector instructions, or the same hardware blocks may be reused.

[0051] As previously mentioned, one type of vector load / store instruction that may be supported is a vector aggregate / scatter instruction. Such a vector instruction may indicate a number of discrete addresses in memory and control the processing circuitry 4 to load data from those discrete addresses into respective elements of a vector register (in the case of a vector aggregate instruction) or store data from those respective elements of a vector register into discrete addresses (in the case of a vector scatter instruction). In accordance with the techniques described herein, rather than using a vector of standard address designations to identify the various memory addresses, a new form of vector aggregate / scatter instruction is provided that may specify a vector of capabilities that are used to determine the various memory addresses. This may provide finer control over the performance of the individual memory access operations used to implement the vector aggregate / scatter operation, since separate capabilities may be defined for use in association with each of those individual memory access operations. In addition to providing an address designation, each capability typically includes constraint information that is used to restrict the operations that may be performed when using that capability. For example, the constraint information may identify a non-extensible range of memory addresses accessible by the processing circuit when using the addressing provided by the capability, and may also provide one or more permission flags identifying associated permissions (e.g., whether read access is allowed, whether write access is allowed, whether access from a specified privilege or security level is allowed, whether the capability can be used to generate memory addresses for instructions to be fetched and executed).

[0052] When executing this new form of vector aggregate / scatter instruction, each data element that is moved between memory and the vector registers (the direction of the move depends on whether a vector aggregate or vector scatter operation is being performed) has an associated capability, and capability access check circuitry 16 within processing circuitry 4 may be used to perform a capability check for each data element to determine whether the memory access operation used to access that given data element is permitted given the constraint information specified by the associated capability. Thus, this may involve checking both whether the memory address is accessible given any range constraint information in the capability, and whether the type of access is permitted given the constraint information in the capability. Further details of how the capabilities required when executing such vector aggregate / scatter instructions are arranged within a series of vector registers are described in more detail with reference to some of the remaining figures.

[0053] As shown in Figure 1, beat control circuitry 20 may be provided to control the operation of the instruction decoder 6 and processing circuitry 4, as appropriate. In particular, in some exemplary implementations, execution of a vector instruction may be divided into portions called "beats," with each beat corresponding to the processing of a portion of a vector of a given size. As will be described in more detail below with reference to Figures 10 and 11, this may allow overlapping execution of vector instructions, thereby improving performance.

[0054] 2 shows a schematic diagram of how tag bits are used in association with individual data blocks to identify whether they represent capabilities or normal data. Specifically, memory address space 110 stores a series of data blocks 115, typically having a specified size. Purely for illustrative purposes, in this example, it is assumed that each data block contains 64 bits, although in other exemplary implementations data blocks of different sizes may be used, for example 128-bit data blocks when a capability is defined by 128 bits of information. In association with each data block 115, in one example, a tag field 120 is provided, which is a single-bit field referred to as a tag bit, which is set to identify that the associated data block represents a capability and is cleared to indicate that the associated data block represents normal data and therefore cannot be treated as a capability. It will be appreciated that the actual values ​​associated with a set or clear state may vary depending on the exemplary implementation, but purely for purposes of illustration, in one exemplary implementation, when the tag bit has a value of 1 it indicates that the associated data block is a capability, and when it has a value of 0 it indicates that the associated data block contains normal data. In one exemplary implementation, the tag bit may not form part of the normal memory address space, but instead may be stored "out of band", for example in a separate tag memory.

[0055] When a capability is loaded into a register 100 accessible to the processing circuit, the tag bits travel with the capability information. Thus, when a capability is loaded into a register 100, an address indication 102 (sometimes referred to herein as a pointer) and metadata 104 providing constraint information (such as the range and permission information discussed above) are loaded into the register. Additionally, in association with the register, or as a particular bit field therein, a tag bit 106 will be set to identify that the contents represent a valid capability. Similarly, when a valid capability is stored back from memory, the associated tag bit 120 will be set in association with the data block in which the capability is stored. Such an approach ensures that capabilities are distinguished from normal data, and therefore normal data cannot be used as a capability.

[0056] The device may be provided with dedicated capability registers (not shown in FIG. 1) for storing capabilities, and thus the register 100 in FIG. 2 may be a dedicated capability register. However, to execute the new type of vector gather / scatter instructions described above, it is desirable to place the required capabilities in several vector registers in the set of vector registers 12. To allow for a distinction between valid capabilities and general-purpose data stored in the vector registers, the set of vector registers is complemented by providing an associated valid capability indication storage, and two different ways in which this may be implemented are shown diagrammatically in FIGS. 3A and 3B. In the example shown in FIG. 3A, the set of vector registers 130 comprises a number of vector registers 135, each of which is of sufficient size to provide several capability size blocks 137. Purely by way of example, when the capabilities are 64 bits long, each capability size block 137 may be 64 bits and the length of each vector register may be 2. N ×64 bits, where N is an integer equal to or greater than 0.

[0057] In the particular example of FIG. 3A, each vector register is assumed to be 128 bits in length, and therefore each vector register has two capability size blocks 137. A valid capability indication storage 140 is provided in association with the set of vector registers, the valid capability indication storage 140 having an entry 145 for each vector register 135. Each entry 145 provides a valid capability indication field for each capability size block 137 in the associated vector register 135. The valid capability indication field may take a variety of forms, but in one exemplary implementation may be a single bit field, and thus in one example may take the form of the tag bit discussed above. In such a case, it will be appreciated that each entry 145 provides a tag bit for each capability size block 137 in the associated vector register 135 to identify whether or not that capability size block stores a valid capability.

[0058] In the example of Fig. 3A, the valid capability indication storage 140 is considered to be a separate structure from the set of vector registers 130, but in alternative implementations, the valid capability indication storage can be effectively incorporated within the set of vector registers by increasing the size of the vector registers to accommodate the required tag bits. In the configuration as shown in Fig. 3B, the set of vector registers 150 includes several capability size blocks 160, 164, each of which has an associated valid capability indication field 162, 166 for storing the associated tag bits. Note that in this configuration, the size of the capabilities is not considered to change, and thus in the example above, each capability is still 64 bits long. However, the vector registers are expanded to provide space for the associated tag bits. 3B, where it is again considered that two capabilities may be stored in each vector register, and assuming each capability is 64 bits long, any vector register that can store capabilities may be configured to be 130 bits long to allow both the two capabilities and their associated tag bits to be stored. In this example, the tag bits are part of the vector register 155, but access to the tag bits may still be tightly controlled, as previously discussed, so that the tag bits are not directly accessible to general purpose processing instructions, and changing a value in a vector register using a non-capability instruction causes the tag to be cleared.

[0059] It should be noted that while in the examples of Figures 3A and 3B it is assumed that all of the vector registers are capable of storing capabilities, in alternative implementations a subset of the vector registers in the set may be reserved for storing capabilities, in which case only the subset of vector registers need be provided with associated valid capability indication storage, whether as separate storage (as in the example of Figure 3A) or incorporated within the vector register structure itself (as in the example of Figure 3B).

[0060] 4A and 4B are flow diagrams illustrating how tag bits maintained in association with each capability size block of a vector register may be managed according to one exemplary implementation. Figure 4A illustrates several steps that are performed to determine what action should be taken in association with the associated tag bits maintained for a capability size block in a vector register being written. Specifically, if at step 170 it is determined that a write operation is being performed to a vector register, then the remainder of the process of Figure 4A is performed for each capability size block in that vector register being written.

[0061] In step 172, it is determined whether the data being written for a given capability size portion of the vector register is the full capability block size. If not, the tag bit is cleared if it was previously set, and the process proceeds to step 174 where the tag bit is cleared. Such an approach prevents illegal modification of capabilities. For example, if an attempt is made to modify a certain number of bits of a valid capability stored in a vector register, the above process causes the tag bit to be cleared, preventing the modified version currently stored in the vector register from being used as a capability.

[0062] However, assuming that a complete capability sized block of information has been written to a given capability sized portion of a vector register, then in step 176 it is determined whether a valid capability has been written. If not, the process proceeds to step 174 where the tag bit is cleared. However, if a valid capability has been written, the process proceeds to step 178 where the tag bit is set.

[0063] It should be noted that tag bits associated with capability-sized blocks in a vector register may not only be cleared during execution of an instruction that writes to the vector register. In particular, as shown by FIG. 4B, in step 180, it may be determined whether any steps have been taken to cause the capabilities stored in the capability-sized blocks of the vector register to no longer be valid. If no such condition is detected, no updates to the associated tag bits are made, as shown by step 185, but whenever the condition is detected, the associated tag bits are cleared in step 190.

[0064] 5A illustrates, in accordance with one exemplary implementation, fields that may be provided within a vector memory access instruction 200 (also referred to herein as a vector gather instruction or a vector scatter instruction). The opcode field 205 may be used to identify the type of vector memory access instruction, and thus, in this example, whether a gather or scatter variant is specified, and that the instruction is of the aforementioned type that uses a capability to determine the memory address to be accessed.

[0065] The data vector designation field 210 is used to identify at least one vector register associated with data elements that are moved between the vector register set and memory through execution of the instruction. In one exemplary implementation, a single vector register is identified by the data vector designation field 210. It will be appreciated that such an identified vector register serves as a source vector register when performing a vector distribution operation, or serves as a destination vector register when performing a vector aggregation operation.

[0066] Also, at least one capability vector indication field 215 may be provided, the contents of which are used to identify multiple vector registers that store the capabilities required to determine the memory addresses of each of the data elements that are to undergo the vector scatter or vector aggregation operation. In one implementation, multiple capability vector indication fields may be provided, e.g., one field for each of the vector registers that contain the required capabilities, while in another exemplary implementation, a single capability vector indication field is used to provide sufficient information to determine which of the vector registers stores the capabilities, and the other vector registers are then determined based on some predetermined relationship. This latter approach may be advantageous from an instruction encoding perspective. The predetermined relationship may take various forms. For example, the vector registers may be contiguous to one another, may form an even / odd pair, or a known offset may exist between the various vector registers.

[0067] As shown in FIG. 5A, the instruction 200 may also include one or more optional fields 220 that capture additional information. For example, an immediate value may be specified that indicates an address offset that may be used in various ways. For example, the address offset may be combined with (e.g., added to) an address instruction in each capability to identify a memory address to be accessed. As another example, the address offset may be used to update an address instruction in each capability (again, for example, by combining the address offset with an existing address instruction) such that the updated capability in the vector register is ready to be used in connection with a subsequent vector memory access instruction. Indeed, in one exemplary implementation, both of the above address instruction adjustment processes may be performed, and typically the same immediate value is used for both adjustment processes.

[0068] As another example of optional information that may be provided in one or more fields 220, information may be provided that specifies the data element size and / or capability size of data elements accessed during execution of the instruction. In some implementations, this information may not be necessary as the capability size may be fixed, and it may be the case that vector memory access instructions of the type described herein are only permitted to be performed on data elements of a particular size, and thus in that illustrative case both the data element size and the capability size are known without having to be separately specified by the instruction.

[0069] Note that while in Figure 5A the various bits forming each field are shown consecutively, this is purely for illustrative purposes and which bits in the instruction are associated with which fields will vary depending on the implementation. Purely by way of example, if the vector register identifier field is four bits wide, three bits will be grouped together, although a fourth bit could be provided elsewhere in the instruction encoding.

[0070] Figure 5B is a flow diagram illustrating the steps performed when executing a vector memory access instruction such as that shown in Figure 5 A. In step 230, it is determined whether the vector memory access instruction is to be executed, and if so, the process proceeds to step 235 where the vector register associated with the data element is determined from the information in the data vector designation field.

[0071] A number of vector registers containing the required capabilities are also determined using information in the at least one capability vector designation field, step 240. As explained above, multiple capability vector designation fields may be provided, each identifying, for example, one of the vector registers, or alternatively, a single capability vector designation field may be provided to allow for the determination of one of the vector registers, with the other vector registers then being determined taking into account known relationships.

[0072] In step 245, for each given data element to which the vector memory access instruction pertains, a memory address of the given data element is determined based on the address indication provided by the associated capability. Furthermore, it is determined whether the memory access operation used to access the given data element is permitted based on the constraint information of the associated capability. This may include not only determining whether the memory address is within an allowable range specified by the range constraint information in the associated capability, but also determining whether any other constraints specified by the metadata of the associated capability are satisfied (e.g., if a vector distribution operation is being performed and therefore the individual memory access operation being performed on the given data element is a write operation, then whether a write access is permitted using the associated capability).

[0073] At step 250, a memory access operation may be allowed to execute for each data element for which it is determined that the memory access operation is permitted. In one exemplary implementation, a memory access operation may be executed for each data element for which the memory access operation is permitted, while in other implementations, it may be determined to inhibit execution of one or more permitted memory access operations if another of the memory access operations is not permitted. As previously mentioned, which permissible access is inhibited in such a situation may depend on the implementation and may depend on where in the vector of data elements is a data element with an associated access that is not permitted.

[0074] 6A is a flow diagram illustrating a technique that may be used to determine vector registers that hold the required capabilities in an implementation in which a single capability vector indication field is provided. In step 300, one vector register that holds the required capabilities is determined from the information in the single capability vector indication field. Then, in step 310, each other vector register that holds the required capabilities is determined from the vector registers identified in step 300 and a known, determined relationship. The determined relationship may be implicit or may be specified in the capability vector indication field or indeed in another field of the instruction.

[0075] 6B illustrates a particular exemplary implementation that may be used to calculate the various vector registers that hold the required capabilities. In step 320, the number of vector registers that contain the required capabilities is determined, and in this exemplary implementation, such vector registers are 2 NIn some implementations, the number of vector registers required to hold the capabilities is fixed, for example because a given vector memory access instruction is only supported for use with data elements of a particular fixed size, and the capabilities are also fixed-sized. However, alternatively, the number of vector registers can be determined at runtime by an instruction decoder, for example based on data element sizes and capability size information specified by the instruction.

[0076] In step 330, a first vector register number is determined from the information provided in the capability vector designation field, with the least significant N bits of that vector register number being constrained to be a logical 0 value in this implementation. It will be appreciated that in such an implementation, the capability vector designation field need not specify those bits since they may be hardwired to 0.

[0077] Each of the other vector register numbers for the plurality of vector registers containing the required capabilities is determined by manipulating the N least significant bits of the first determined vector register number, step 340. This provides a particularly simple and efficient mechanism for specifying the plurality of vector registers containing the required capabilities.

[0078] FIG. 7 illustrates how a set of vector registers 350 may be considered to be formed from multiple logical sections 360, 365. Each vector register 355 has portions 357, 359 within each section. Although two sections are shown in FIG. 7, in other implementations, three or more sections may be provided. In some implementations, only a single capability is provided per portion 357, 359 of a vector register, while in other implementations, each portion of a register may be large enough to hold multiple capabilities. By such an approach, this may allow the execution of a vector instruction, including a given vector memory access instruction, to be divided into multiple "beats", during which only one section of the set of vector registers is accessed to execute the vector instruction. By allowing a vector instruction to be divided into multiple beats, it may be possible for the execution of a vector instruction to overlap with the execution of one or more other vector instructions, resulting in a very efficient implementation. For example, a given vector memory access instruction may overlap with a vector operation instruction. In particular, in one exemplary implementation, all of the data elements and capabilities required to perform memory access operations during any particular beat may be obtained from a single section of the set of vector registers, then leaving any other sections available for access during execution of the overlapped instructions. Further details of a beat-based implementation are described in more detail below with reference to Figures 10 and 11.

[0079] 8A-8C illustrate different specific exemplary configurations of data elements and associated capabilities that may be used when performing a gather or scatter operation as described herein. As shown in FIG. 8A, the term "CX" identifies a capability used to determine a memory address for a corresponding data value "DX". The vector register 400 shown in FIG. 8A is a vector register associated with data elements accessed during execution of a vector memory access instruction. In this exemplary implementation, assume that the vector register 400 is 128 bits wide and each data element is 32 bits wide, such that four data elements are associated with the vector register 400. Each data element may be considered to be associated with a corresponding data lane of the vector register 400, and thus, as shown in FIG. 8A, the data lanes may take on values ​​0-3.

[0080] In the example shown in Figures 8A-8C, the capabilities are 64 bits wide, and thus each vector register 405, 410 in the example of Figure 8A can store two capabilities (for purposes of illustration in Figures 8A-8C, any additional bits provided to hold the aforementioned tag values ​​are omitted). Each capability in a particular vector register can be considered to occupy an associated capability lane, and thus in the example of Figure 8A, there are two capability lanes, referred to as lanes 0 and 1. As shown in Figure 8A, capability C0 is stored in the first capability register Q. N 405, occupies capability lane 0, and capability C1 is stored in the second capability register Q N+1 410, and capability C2 is stored in the first capability register Q N 405, occupies capability lane 1, and capability C3 is stored in the second capability register Q N+1410. Thus, with such a configuration, it can be appreciated that for any given pair of data elements associated with adjacent locations in vector register 400, the associated capabilities are stored in different ones of the plurality of vector registers 405, 410.

[0081] Such an arrangement has been found to be highly advantageous as it means that the capabilities required in relation to a particular sequence of data elements can all be found in the same parts 357, 359 of the vector register. In particular, in the example shown in Figure 8A, data elements D0 and D1 and the capabilities C0 and C1 required to identify the memory addresses of those data elements can all be found in the bottom half of the associated vector register, and similarly data elements D2 and D3 and the capabilities C2 and C3 required to identify the memory addresses of those data elements can all be found in the top half of the associated vector register. This can, for example, support beat-by-beat execution of the vector memory access instructions referred to above.

[0082] In Figure 8A, the data values ​​are 32 bits, but this is not a requirement and Figure 8B shows an alternative example where the data elements are 16 bits wide. Thus, a 128-bit wide vector register 415 can be associated with eight data elements, and four vector registers Q N ~Q N+3 Vector registers 420, 425, 430, 435 are required to hold the associated capabilities. Again, the capabilities are laid out similarly to Figure 8A, with the first four capabilities stored in the bottom halves of vector registers 420, 425, 430, 435 and the last four capabilities stored in the top halves of those vector registers.

[0083] It is also not a requirement that the vector registers are considered to be 128-bit registers; in the example of FIG. 8C, each of the registers is 256-bits wide. In this particular example, the data elements are 32-bits wide, and the capabilities remain the same as in the other examples, i.e., 64-bits wide. In this example, it can thus be seen that there are eight data elements associated with the vector register 440, and two vector registers 445, 450 are used to store eight capabilities, with four capabilities located within each register. The capabilities are organized such that they are stored in ascending order in capability lane 0, capability lane 1, capability lane 2, and capability lane 3, thus following the general pattern previously described with reference to the other two examples of FIG. 8A and FIG. 8B.

[0084] When performing the beat-by-beat execution of the vector memory access instructions described above, in one exemplary implementation, each section of the vector register may be configured to store one or more capabilities. Thus, considering the example of FIG. 8A or FIG. 8B, the vector register may be considered to be formed of two sections, allowing half of the required access operations to be processed in the first beat and the other half to be processed in the second beat. Similarly, considering FIG. 8C, the vector register set may be considered to be formed of two or four sections, allowing the required access operations to be executed over two or four beats, respectively. However, it should be noted that it is not necessarily a requirement that each logical section of the vector register be wide enough to accommodate at least one capability. For example, in some implementations, it may be possible to have a section size smaller than the capability size, for example, a 32-bit section size with a 64-bit capability.

[0085] Figure 9 is a flow diagram showing how an associated capability may be determined for each data element when using a capability layout as shown diagrammatically in Figures 8A-8C. In step 450, a parameter M is set equal to the number of capability lanes and a parameter P is set equal to the number of vector registers that hold the capability. In step 455, the vector registers are considered to be identified by a sequence of values ​​0 to P-1 and the capability lanes are considered to be identified by a sequence of values ​​0 to M-1. In step 460, a parameter X is set to 0 and then in step 465, a calculation X / P is performed on the data elements in lane X.

[0086] In step 470, the quotient and remainder from the above calculations are used to identify the capability lane and vector register, respectively, that contain the associated capability. In step 475, it is determined whether data lane X is the last data lane, and if not, the value of X is incremented in step 480 before returning to step 465. If in step 475, it is determined that data lane X is the last data lane, then the process ends in step 485.

[0087] In some applications, such as digital signal processing (DSP), there may be an approximately equal number of ALU and load / store instructions, and thus some large blocks, such as the MAC, may be left idle for a significant amount of time. This inefficiency may be exacerbated on vector architectures, as execution resources scale with the number of vector lanes to obtain higher performance. On smaller processors (e.g., single-issue, in-order cores), the area overhead of a fully scaled-out vector pipeline may be very large. One technique to minimize the area impact while making better use of the available execution resources is to overlap instruction execution, as shown in FIG. 10. In this example, three vector instructions include a load instruction VLDR, a multiply instruction VMUL, and a shift instruction VSHR, all of which can execute simultaneously, even if there is a data dependency between them. This is because element 1 of VMUL depends only on element 1 of Q1, and not on the entire Q1 register, and therefore execution of VMUL can begin before VLDR finishes executing. By overlapping instructions, expensive blocks, such as multipliers, can be kept active for more time.

[0088] Therefore, it may be desirable to allow a microarchitectural implementation to overlap the execution of vector instructions. However, if an architecture assumes that there is a fixed amount of instruction overlap, then the microarchitectural implementation may provide high efficiency if it actually matches the amount of instruction overlap assumed by the architecture, but may cause problems when scaled to a different microarchitecture that uses a different overlap or no overlap at all.

[0089] Instead, the architecture may support a range of different overlaps, as shown in the example of FIG. 11. The execution of a vector instruction is divided into parts called "beats", each beat corresponding to the processing of a portion of a vector of a given size. A beat is an infinitesimal portion of a vector instruction that is either fully executed or not executed at all; it cannot be executed partially. The size of the portion of the vector processed in one beat is defined by the architecture and can be any fraction of the vector. In the example of FIG. 11, a beat is defined as the processing corresponding to one-quarter of the vector width, so there are four beats per vector instruction. Obviously, this is only an example, and other architectures may use a different number of beats, e.g., two or eight. The portion of the vector corresponding to one beat may be the same size as the data element size of the vector being processed, or it may be larger or smaller. Thus, a beat is a specific fixed width of vector processing, even if the element size varies from implementation to implementation or at execution time between different instructions. If the portion of the vector being processed in one beat contains multiple data elements, the carry signal can be disabled at the boundary between each element to ensure that each element is processed independently. If the portion of the vector processed in one beat corresponds to only a portion of the elements and the hardware is insufficient to compute several beats in parallel, the carry output generated during one beat of processing may be input as the carry input to the next beat of processing, such that the results of the two beats together form a data element.

[0090] As shown in FIG. 11, different microarchitectural implementations of processing circuitry 4 may execute different numbers of beats in one "tick" of the abstract architectural clock, where a "tick" corresponds to a unit of progression of the architectural state (e.g., in a simple architecture, each tick may correspond to an instance of updating all architectural state associated with the execution of an instruction, including updating a program counter to point to the next instruction). It will be understood by those skilled in the art that known microarchitectural techniques such as pipelining may mean that a single tick may require multiple clock cycles to execute at the hardware level, and in fact, that a single clock cycle at the hardware level may process multiple portions of multiple instructions. However, such microarchitectural techniques are invisible to software, since ticks are atomic at the architectural level. For the sake of brevity, such microarchitectures will be ignored during further description of this disclosure.

[0091] As shown in the lower example of Figure 11, some implementations may schedule all four beats of a vector instruction in the same tick by providing sufficient hardware resources to process all beats in parallel within one tick. This may be suitable for higher performance implementations, where no overlap between instructions at the architectural level is required since the entire instruction can be completed in one tick.

[0092] On the other hand, a more area-efficient implementation may provide a narrower processing unit that can only process two beats per tick, and as shown in the central example of FIG. 11, instruction execution may overlap with the first and second beats of a second vector instruction executing in parallel with the third or fourth beat of the first instruction, and these instructions executing on different execution units within the processing circuit (e.g., in FIG. 11, the first instruction is a load instruction executed using load / store unit 18 (which may be, for example, a vector aggregation instruction of the type described herein), and the second instruction is a multiply-accumulate instruction executed using a MAC unit provided within processing circuit 4).

[0093] An even more energy / area efficient implementation could provide a hardware unit that is narrower and can only process a single beat at a time, where one beat can be processed per tick and instruction execution is overlapped and staggered by two beats as shown in the top example of Figure 11. In one exemplary implementation, the section size can be used to affect the amount of staggering between instructions (as it is desirable to get all of the data from the same section when executing a particular beat). In the top example shown in Figure 11, for example, the beat size is 32 bits, but the section size may be 64 bits, and thus this is why the instructions are staggered by two beats.

[0094] It will be appreciated that the overlaps shown in Figure 11 are just some examples and other implementations are possible. For example, some implementations of processing circuitry 4 can support dual issue of multiple instructions in parallel in the same tick, thereby increasing instruction throughput. In this case, two or more vector instructions that start together in one cycle may have some beats of overlap with two or more vector instructions that start in the next cycle.

[0095] In addition to varying the amount of overlap from implementation to scale to different performance points, the amount of overlap between vector instructions may also vary at run-time between different instances of execution of vector instructions in a program. Thus, processing circuitry 4 may include beat control circuitry 20 as shown in FIG. 1 to control when a given instruction is executed relative to a previous instruction. This gives the microarchitecture the freedom to choose not to overlap instructions in certain tricky cases that are more difficult to implement, or depending on the resources available to the instructions. For example, if there are consecutive instructions of a given type (e.g., multiply-accumulate) that require the same resource, and all available MAC or ALU resources are already in use by another instruction, there may not be enough free resources to start executing the next instruction, and therefore, rather than overlapping, issue of the second instruction may wait until the first instruction is completed.

[0096] FIG. 12 is a flow diagram showing how a sequence of vector capability memory transfer instructions can be used to move a set of capabilities between memory and a number of vector registers, while performing the necessary reconfiguration to ensure that those capabilities are stored in the vector registers, in the form of a configuration shown by way of example with reference to the previous examples of FIGS. 8A-8C.

[0097] At step 490, a sequence of vector capability memory transfer instructions is decoded, with each such instruction defining an associated access pattern and identifying the subset of capabilities required by any particular instance of the aforementioned vector aggregate / scatter instruction. In one exemplary implementation, each individual vector capability memory transfer instruction identifies a different subset of capabilities relative to each other vector capability memory transfer instruction in the sequence.

[0098] Capabilities are then moved between memory and the identified vector registers, with de-interleaving (if a load operation is being performed) or interleaving (if a store operation is being performed) as defined by the access pattern of each vector capability memory transfer instruction, in step 492. As a result, capabilities may be configured to be stored sequentially in memory, while in vector registers, the capabilities are de-interleaved such that any given pair of capabilities stored sequentially in memory are stored in different vector registers.

[0099] The vector capability memory transfer instructions used to perform the steps shown in Figure 12 do not have to directly follow each other in program order and therefore do not have to be executed one after the other sequentially. Once all the vector capability memory transfer instructions in a sequence have been executed, the reconfiguration of capabilities required to be moved between memory and vector registers is performed.

[0100] In one exemplary implementation, the memory is formed from multiple memory banks, and for each vector capability memory transfer instruction, an access pattern is defined such that two or more of the memory banks are accessed when the vector capability memory transfer instruction is executed. Banked memories make it easier for hardware to implement parallel transfers to / from the memory, and therefore it is beneficial to specify an access pattern that allows this. This is shown diagrammatically in FIG. 13 as an example of a memory formed from two memory banks 496, 498, each memory bank being 64 bits wide. In such an arrangement of memory banks, when the memory access logic 494 is processing a memory address, it can take into account bit 3 of the address to determine which bank to access. In particular, if bit 3 of the address (i.e., the fourth address bit, assuming the first address bit is bit 0) is a logic 0 value, the memory bank 496 is accessed, and if bit 3 of the address is a logic 1 value, the other memory bank 498 is accessed. It will be appreciated that since the capabilities are 64-bit capabilities, odd capabilities are stored in one bank and even capabilities are stored in the other bank.

[0101] Purely by way of example, considering the capability configuration shown in FIG. 8A, capabilities C0-C3 arranged sequentially in memory may be loaded into capability registers 405 and 410 using two vector capability memory transfer instructions as follows: VLDRC2_1:C0→Qn[63:0], C3→Q(n+1)[127:64] VLDRC2_2:C1→Q(n+1)[63:0], C2→Qn[127:64]

[0102] 12, it can be seen that capability C0 is in a different bank than capability C3, capability C1 is in a different bank than capability C2, and so both banks 496, 498 are accessed when executing each of these instructions. Also, the two capabilities transferred by each instruction reside in different capability lanes of the vector register and therefore, in one embodiment, can write to the vector register at the same time.

[0103] FIG. 14 illustrates a simulator implementation that may be used. While the above examples implement the invention in terms of apparatus and methods for operating specific processing hardware that supports the technique, it is also possible to provide an instruction execution environment according to the examples described herein, which is implemented through the use of a computer program. Such computer programs are often referred to as simulators insofar as they provide a software-based implementation of a hardware architecture. Various simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on a host processor 515, which optionally runs a host operating system 510 and supports the simulator program 505. In some arrangements, there may be multiple layers of simulation between the hardware and the provided instruction execution environment, and / or multiple different instruction execution environments may be provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations that run at reasonable speeds, but such an approach may be justified in certain situations, such as when it is desirable to run code native to another processor for compatibility or reuse reasons. For example, a simulator implementation may provide an instruction execution environment that has additional functionality not supported by the host processor hardware, or that is typically associated with a different hardware architecture. An overview of simulation is given in "Some Efficient Architecture Simulation Techniques", Robert Bedichek, Winter 1990 USENIX Conference, pp. 53-63.

[0104] To the extent that examples have been described above with reference to particular hardware constructs or features, in the simulated implementations, equivalent functionality may be provided by suitable software constructs or features. For example, particular circuits may be provided as computer program logic in the simulated implementations. Similarly, memory hardware such as registers or caches may be provided as software data structures in the simulated implementations. Also, the physical address space used to access memory 8 in hardware device 2 may be emulated as a simulated address space, which is mapped by simulator 505 to a virtual address space used by host operating system 510. In arrangements where one or more of the hardware elements referred to in the preceding examples are present in host hardware (e.g., host processor 515), some simulated implementations may use the host hardware, if suitable.

[0105] The simulator program 505 may be stored in a computer-readable storage medium (which may be a non-transitory medium) and provides a virtual hardware interface (instruction execution environment) to the target code 500 (which may include applications, operating systems, and hypervisors), the virtual hardware interface being the same as the hardware interface of the hardware architecture modeled by the simulator program 505. Thus, the program instructions of the target code 500 may be executed from within the instruction execution environment using the simulator program 505, so that the host computer 515, which does not actually have the hardware features of the device 2 discussed above, can emulate these features. The simulator program may include processing program logic 520 that emulates the operation of the processing circuit 4, instruction decode program logic 525 that emulates the operation of the instruction decoder 6, and vector register emulation program logic 522 that maintains a data structure to emulate the vector register 12. Thus, the techniques described herein for performing vector aggregation or scatter operations using capabilities may be implemented in software by the simulator program 505 in the example of FIG. 14.

[0106] In this application, the term "configured to..." is used to mean that an element of an apparatus has a configuration that is capable of performing a defined operation. In this context, "configuration" refers to a manner of arrangement or interconnection of hardware or software. For example, an apparatus may have dedicated hardware that provides the defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not imply that an apparatus element needs to be modified in any way to provide the defined operation.

[0107] Although illustrative examples of the invention have been described in detail herein with reference to the accompanying drawings, it will be understood that the invention is not limited to these precise examples, and that various changes, additions and modifications may be made to these examples by those skilled in the art without departing from the scope of the invention as defined by the appended claims. For example, various combinations of the features of the dependent claims may be made with the features of the independent claims without departing from the scope of the invention.

Claims

1. 1. An apparatus comprising: processing circuitry for performing vector processing operations; a set of vector registers; an instruction decoder that decodes a vector instruction and controls the processing circuitry to perform the vector processing operation specified by the vector instruction; the instruction decoder, in response to a given vector memory access instruction specifying a plurality of memory access operations, each memory access operation being performed to access an associated data element, determines, from a data vector designation field of the given vector memory access instruction, at least one vector register in the set of vector registers associated with a plurality of data elements; and determines, from at least one capability vector designation field of the given vector memory access instruction, a plurality of vector registers in the set of vector registers including a plurality of capabilities, each capability being associated with one of the data elements in the plurality of data elements and providing an address designation and constraint information constraining use of the address designation when accessing memory; the number of vector registers determined from the at least one capability vector designation field being greater than the number of vector registers determined from the data vector designation field; The instruction decoder configures the processing circuitry to: for each given data element within the plurality of data elements, determining a memory address based on the address indication provided by the associated capability, and determining whether the memory access operation used to access the given data element is permitted with respect to the determined memory address taking into account the constraint information of the associated capability; and controlling the execution of the memory access operation for each data element for which the memory access operation is permitted, such that execution of the memory access operation for any given data element moves the given data element between the determined memory address in the memory and the at least one vector register.

2. 2. The apparatus of claim 1, further comprising: capability indication storage providing a valid capability indication field associated with each capability-sized block in a given vector register of the set of vector registers, each valid capability indication field configured to be set to indicate when the associated capability-sized block stores a valid capability and is otherwise cleared.

3. The apparatus of claim 2 , wherein the capability indication storage is embedded within the set of vector registers.

4. 4. The apparatus of claim 2 or 3, wherein the processing circuitry is configured to allow only any valid capability indication field to be set to indicate that a valid capability is stored in the associated capability-sized block in response to execution of one or more particular instructions in a set of instructions executable by the apparatus.

5. 4. The apparatus of claim 1, wherein the number of vector registers forming the plurality of vector registers determined from the at least one capability vector indication field is a power of two.

6. 4. The apparatus of claim 1, wherein the at least one capability vector designation field is a single capability designation field configured to identify one vector register, and the instruction decoder is configured to determine remaining vector registers of the plurality of vector registers based on the determined relationship.

7. The number of vector registers in the plurality of vector registers is 2 N 7. The apparatus of claim 6, wherein a single capability vector designation field indicates a first vector register number identifying the one vector register, the first vector register number being constrained to have its N least significant bits set to logic 0 values, and the instruction decoder is configured to generate vector register numbers for each of the remaining vector registers by reusing the first vector register number and selectively setting at least one of the N least significant bits to a logic 1 value.

8. 4. The apparatus of claim 1, wherein for any given pair of data elements associated with adjacent positions in the at least one vector register, the associated capabilities are stored in different vector registers of the plurality of vector registers.

9. the at least one vector register determined from the data vector indication field comprises a single vector register, each data element being associated with a corresponding data lane of the single vector register; each capability is located in a capability lane in one of the vector registers in the plurality of vector registers; 4. The apparatus of claim 1, wherein for a given data element, the vector register containing the associated capability is determined according to a given number of least significant bits of a lane number of the corresponding data lane, and the capability lane containing the associated capability is determined according to the remaining bits of the lane number of the corresponding data lane.

10. the number of vector registers in the plurality of vector registers containing the plurality of capabilities is P, logically viewed as a sequence having values ​​0 to P-1, and the number of capability lanes in any given vector register is M, having values ​​0 to M-1; 10. The apparatus of claim 9, wherein the data lane associated with the given data element is data lane X having values ​​0 to X-1, and the location of the associated capability within the plurality of vector registers is determined by dividing X by P to produce a quotient and a remainder, the quotient identifying the capability lane containing the associated capability, and the remainder identifying the vector register containing the associated capability.

11. the set of vector registers is logically divided into a plurality of sections, each section containing a corresponding portion from each of the vector registers in the set of vector registers; the plurality of capabilities are located within the plurality of vector registers such that, for each data element, the associated capability is stored within the same section as the data element; 4. The apparatus of claim 1, wherein execution of the given vector memory access instruction is divided into multiple beats, and during each beat, only one section of the set of vector registers is accessed to execute the given vector memory access instruction.

12. 12. The apparatus of claim 11, wherein the processing circuitry is configured to perform the memory access operations on the data elements in a given section over one or more beats before performing the memory access operations on the data elements in a next section over one or more beats.

13. 4. The apparatus of claim 1, wherein the instruction decoder is configured to decode a plurality of vector capability memory transfer instructions that cause the instruction decoder to control the processing circuit to transfer a plurality of capabilities between the memory and the plurality of vector registers, and to reconfigure the plurality of capabilities during the transfer such that the plurality of capabilities are stored sequentially in the plurality of vector registers and deinterleaved across the plurality of vector registers such that any given pair of capabilities within the plurality of capabilities stored sequentially in the memory are stored in different ones of the plurality of vector registers.

14. 14. The apparatus of claim 13, wherein each vector capability memory transfer instruction is configured to identify a different capability relative to each other vector capability memory transfer instruction, and wherein each vector capability memory transfer instruction is configured to identify an access pattern that causes the processing circuitry to transfer the identified capability while performing the reconfiguration specified by the access pattern.

15. the memory is composed of a plurality of memory banks, 15. The apparatus of claim 14, wherein for each vector capability memory transfer instruction, the access pattern is defined such that two or more of the memory banks are accessed when the vector capability memory transfer instruction is executed by the processing circuitry.

16. 4. The apparatus of claim 1, wherein the at least one vector register determined from the data vector designation field of the given vector memory access instruction comprises a single vector register, the capability is twice the size of the data element, and the plurality of vector registers determined from the at least one capability vector designation field comprises two vector registers.

17. 4. The apparatus of claim 1, wherein the given vector memory access instruction further includes an immediate value indicating an address offset, and wherein the processing circuitry is configured to, for each given data element in the plurality of data elements, determine the memory address of the given data element by combining the address offset with the address indication provided by the associated capability.

18. 4. The apparatus of claim 1, wherein the given vector memory access instruction further includes an immediate value indicating an address offset, and wherein, for each given data element, the processing circuitry is configured to update the address indications of the associated capabilities in the plurality of vector registers by adjusting the address indications according to the address offset.

19. 1. A method of performing memory access operations in an apparatus providing processing circuitry for performing vector processing operations and a set of vector registers, comprising: an instruction decoder responsive to a given vector memory access instruction specifying a plurality of memory access operations, each memory access operation being performed to access an associated data element, determining from a data vector designation field of the given vector memory access instruction at least one vector register in the set of vector registers associated with a plurality of data elements, and determining from at least one capability vector designation field of the given vector memory access instruction a plurality of vector registers in the set of vector registers including a plurality of capabilities, each capability being associated with one of the data elements in the plurality of data elements and providing an address designation and constraint information constraining use of the address designation when accessing memory, the number of vector registers determined from the at least one capability vector designation field being greater than the number of vector registers determined from the data vector designation field; The processing circuitry for each given data element within the plurality of data elements, determining a memory address based on the address indication provided by the associated capability, and determining whether the memory access operation used to access the given data element is permitted with respect to the determined memory address taking into account the constraint information of the associated capability; and controlling such that enabling the memory access operation for each data element for which the memory access operation is permitted, and execution of the memory access operation for any given data element moves the given data element between the determined memory address in the memory and the at least one vector register.

20. 1. A computer program for controlling a host data processing device to provide an instruction execution environment, comprising: processing program logic for performing vector processing operations; vector register emulation program logic that emulates a set of vector registers; instruction decode program logic that decodes vector instructions and controls the processing program logic to perform the vector processing operations specified by the vector instructions; the instruction decode program logic is responsive to a given vector memory access instruction specifying a plurality of memory access operations, each memory access operation being performed to access an associated data element, to determine from a data vector designation field of the given vector memory access instruction at least one vector register in the set of vector registers associated with a plurality of data elements, and to determine from at least one capability vector designation field of the given vector memory access instruction a plurality of vector registers in the set of vector registers comprising a plurality of capabilities, each capability being associated with one of the data elements in the plurality of data elements and providing an address designation and constraint information constraining use of the address designation when accessing memory, the number of vector registers determined from the at least one capability vector designation field being greater than the number of vector registers determined from the data vector designation field; The instruction decode program logic executes the processing program logic as follows: for each given data element within the plurality of data elements, determining a memory address based on the address indication provided by the associated capability, and determining whether the memory access operation used to access the given data element is permitted with respect to the determined memory address taking into account the constraint information of the associated capability; 10. The computer program product of claim 9, further configured to control: enabling execution of the memory access operation for each data element for which the memory access operation is permitted; and controlling such that execution of the memory access operation for any given data element moves the given data element between the determined memory address in the memory and the at least one vector register.