Stack pointer switch validity check

By implementing the stack pointer switching effectiveness check operation in the processing circuit, verifying whether the incoming data value meets the conditions for specifying the paging address, the problem that the stack pointer switching may bypass security measures is solved, and effective supervision of stack pointer switching and system security and stability are achieved.

CN120077369APending Publication Date: 2025-05-30ARM LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380074203.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-27
Filing Date
2023-09-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively regulate the switching of stack pointers, especially when fast software thread switching is required, which may bypass security measures and lead to data access errors.

Method used

By implementing the stack pointer switching validity check operation in the processing circuit, verify that the incoming data value meets the validity condition of at least one stack cap value, which includes the condition of specifying the paging address to ensure that the stack pointer switches to the correct stack structure.

Benefits of technology

Effectively supervise the switching of stack pointers to prevent incorrect stack pointer values ​​from causing data access errors, and improve system security and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077369A_ABST
    Figure CN120077369A_ABST
Patent Text Reader

Abstract

Processing circuitry 16 performs a stack pointer switch validity check operation associated with a switch of the stack pointer from an outgoing stack pointer value to an incoming stack pointer value. The validity check operation includes verifying whether an incoming data value obtained by memory access circuitry 26 in response to a memory access request specifying an address determined based on the incoming stack pointer value conforms to at least one stack cap value validity condition, the at least one stack closure value validity condition includes a condition that a predetermined portion of the incoming data value corresponds to a given page address indicating a page of an address space including the address determined based on the incoming stack pointer value. The at least one stack capping value validity condition is determined irrespective of whether another portion of the incoming data value other than the predetermined portion corresponds to a sub-paging address bit of the address determined based on the incoming stack pointer value. The error handling response is triggered in response to determining that the incoming data value does not conform to the at least one stack cap value validity condition.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This technology relates to the field of data processing.

[0002] A stack pointer can be used to manage access to a stack data structure in memory. A stack data structure is a data structure in which access to entries on the stack is managed according to a last-in, first-out (LIFO) policy. The stack pointer provides an address relative to which the address of an item on the stack can be determined. As items are pushed onto and popped from the stack, the stack pointer is updated to keep track of the location representing the "top" of the stack.

[0003] At least some examples provide an apparatus including: a memory access circuit configured to perform stack access to a stack data structure based on a stack pointer; and a processing circuit configured to perform a stack pointer switch validity check operation associated with a switch of the stack pointer from an outgoing stack pointer value to an incoming stack pointer value, the stack pointer switch validity check operation including verifying whether an incoming data value obtained by the memory access circuit in response to a memory access request specifying an address determined based on the incoming stack pointer value conforms to at least one stack cap value validity condition, the at least one stack cap value validity condition including a condition that a predetermined portion of the incoming data value corresponds to a given paging address indicating paging of an address space including the address determined based on the incoming stack pointer value, wherein: the processing circuit is configured to determine whether the incoming data value conforms to the at least one stack cap value validity condition regardless of whether another portion of the incoming data value other than the predetermined portion corresponds to a sub-paging address bit of the address determined based on the incoming stack pointer value; and the processing circuit is configured to trigger an error handling response in response to determining that the incoming data value does not conform to the at least one stack cap value validity condition.

[0004] At least some examples provide a method that includes: performing a stack pointer switch validity check operation associated with a switch of a stack pointer from an outgoing stack pointer value to an incoming stack pointer value, the stack pointer switch validity check operation including verifying whether an incoming data value obtained by a memory access circuit in response to a memory access request specifying an address determined based on the incoming stack pointer value conforms to at least one stack capping value validity condition, the at least one stack capping value validity condition including a condition that a predetermined portion of the incoming data value corresponds to a given paging address indicating paging of an address space including the address determined based on the incoming stack pointer value, wherein: whether the incoming data value conforms to the at least one stack capping value validity condition is determined regardless of whether another portion of the incoming data value other than the predetermined portion corresponds to sub-paging address bits of the address determined based on the incoming stack pointer value; and the method includes triggering an error handling response in response to determining that the incoming data value does not conform to the at least one stack capping value validity condition.

[0005] At least some examples provide a computer program including instructions that, when executed by a host data processing device, control the host data processing device to provide an instruction execution environment for executing target code. The computer program includes: memory access program logic configured to perform stack access to a stack data structure based on a stack pointer; and handler program logic configured to perform a stack pointer switch validity check operation associated with a switch of the stack pointer from an outgoing stack pointer value to an incoming stack pointer value, the stack pointer switch validity check operation including verifying whether an incoming data value obtained by the memory access program logic in response to a request specifying an address determined based on the incoming stack pointer value conforms to at least one stack capping value validity condition, the at least one stack capping value validity condition including a condition that a predetermined portion of the incoming data value corresponds to a given paging address indicating paging of an address space including the address determined based on the incoming stack pointer value, wherein: the handler program logic is configured to determine whether the incoming data value conforms to the at least one stack capping value validity condition regardless of whether another portion of the incoming data value other than the predetermined portion corresponds to sub-paging address bits of the address determined based on the incoming stack pointer value; and the handler program logic is configured to trigger an error handling response in response to determining that the incoming data value does not conform to the at least one stack capping value validity condition.

[0006] Further aspects, features, and advantages of the present technology will be apparent from the following example description read in conjunction with the accompanying drawings, in which:

[0007] Figure 1 An example of a data processing device is shown;

[0008] Figure 2 Shows an example of a function call;

[0009] Figure 3 Shows an example of a protection control stack (GCS) push and pop operation;

[0010] Figure 4 Shows the access permission check for GCS load / store operations;

[0011] Figure 5 Shows the access permission check for non-GCS load / store operations;

[0012] Figure 6 Shows the permission check for instruction fetch or branch operations;

[0013] Figure 7 Shows a method for control stack switching to switch the stack pointer from an outgoing stack pointer value to an incoming stack pointer value;

[0014] Figure 8 Shows the processing of a first stack pointer switching instruction;

[0015] Figure 9 Shows the processing of a second stack pointer switching instruction;

[0016] Figure 10 Shows stack pointer switching;

[0017] Figure 11 Shows an example encoding of GCS records; and

[0018] Figure 12 Shows an example simulation.

[0019] A device includes a memory access circuit that performs stack access to a stack data structure based on a stack pointer. Different parts of software can be associated with different stack structures in the memory, so when switching between those parts of the software, it may be necessary to switch the corresponding stack pointer from an outgoing stack pointer value associated with the outgoing software to an incoming stack pointer value associated with the incoming software. However, for some stacks, it may be important that the stack pointer cannot be switched to any arbitrary location that is not intended to be the incoming stack structure. For example, when pushing items onto a stack and popping items from the stack, some stacks can be associated with certain security measures. If the stack pointer could be switched to any software-selected address, then this could bypass those security measures because it could lead to future stack accesses to data that has not been protected by those measures.

[0020] Accordingly, techniques for regulating the switching of the stack pointer may be required. One approach could be to trap updates to the stack pointer into a higher-privilege execution state such that higher-privilege software can examine the incoming stack pointer value and determine whether the update is allowable. However, fast software thread switching may need to be supported without invoking the kernel or other higher-privilege software. Such thread switches may occur frequently, and thus the performance burden of invoking the kernel at every switch of the stack pointer may be considered too high.

[0021] In the example discussed below, the processing circuitry supports performing a stack pointer switch validity check operation associated with the switch of the stack pointer from an outgoing stack pointer value to an incoming stack pointer value. The stack pointer switch validity check operation includes verifying whether an incoming data value obtained by the memory access circuitry in response to a memory access request specifying an address determined based on the incoming stack pointer value conforms to at least one stack cap value validity condition. The at least one stack cap value validity condition includes a condition that a predetermined portion of the incoming data value corresponds to a given paging address indicating paging of an address space including the address determined based on the incoming stack pointer value. The processing circuitry determines whether the incoming data value conforms to the at least one stack cap value validity condition, regardless of whether another portion of the incoming data value other than the predetermined portion corresponds to a sub-paging address bit of the address determined based on the incoming stack pointer value. An error handling response is triggered in response to determining that the incoming data value does not conform to the at least one stack cap value validity condition.

[0022] Accordingly, the incoming data value is obtained by the memory access circuitry from a location having a memory address determined based on the incoming stack pointer value. For example, if the incoming stack pointer value does represent a previously established stack structure, then the incoming data value may be obtained from a location that should represent the top of the stack. For the validity check to pass, the incoming data value needs to specify, in a predetermined portion, the given paging address that indicates paging of the address space including the address of the location providing the incoming data value itself. Thus, software needs to ensure that the non-active stack has a stack cap value at a location relative to the incoming stack pointer that will later be used to reference the stack, the stack cap value specifying the paging address of the address space in which the value is stored. This provides a security measure for filtering out some erroneous operations that provide an erroneous value for the incoming stack pointer value that does not correspond to the expected stack structure to be switched to, since it is relatively unlikely that the location represented by an erroneous incoming stack pointer value will happen to contain a data value specifying the paging address of that location itself.

[0023] It might be thought that security could be further enhanced by requiring that, in order to be a valid value, an incoming data value should specify the specific address of the location providing that value, including sub-page address bits that distinguish different addresses within the same page. This could further enhance security because pointers to the same address of the location storing the pointer itself would be extremely rare, making the likelihood of accidentally passing a stack pointer switch validity check operation when specifying a random other address as the incoming stack pointer value extremely low.

[0024] However, the inventors have recognized that, in practice, for security it may be sufficient to indicate the page address in the stack capping value that is used to identify the stack structure that can be effectively switched to. By determining whether at least one stack capping value validity condition is satisfied by the incoming data value, regardless of whether any part of the incoming data value corresponds to the sub-page address bits of the address determined from the incoming stack pointer value, this frees up bits of the data word that provides the stack capping value for other purposes. This can be particularly useful because for some usage scenarios it may be necessary to support several different stack record types that can be allocated on the stack. If sub-page address bits were required to be meaningful when checking the validity of the stack capping value, then most of the bits of the data word would be required to specify the address in the stack capping value, in addition to several lower bits of that address (which may not require any data / instruction address alignment restrictions to be specified explicitly), leaving very few bits available for encoding the stack record type. In contrast, by requiring a valid stack capping value to specify the page address of the location storing the stack capping value, but not requiring the specification of sub-page address bits, this frees up a larger number of bits for encoding the stack record type, which can be extremely useful for supporting future expansion of the processor architecture.

[0025] The given page address can be a virtual page address indicating the page of the virtual address space that includes the virtual address determined based on the incoming stack pointer value. In order to be valid upon a switch of the stack pointer, by requiring the incoming data value (obtained via a memory access circuit based on the incoming stack pointer value) to provide the virtual page address of the location storing the incoming data value, this also provides a sanity check that the incoming stack is accessed with the same virtual-to-physical address translation mapping that was used during the previous access to the stack, which can help detect an attack that confuses the virtual-to-physical address translation mapping by mapping different virtual addresses to the physical address of the location storing the stack based on an attacker's definition.

[0026] The processing circuit can support performing an outgoing stack capping operation associated with the switch of the stack pointer from the outgoing stack pointer value to the incoming stack pointer value, so as to push a valid stack capping value into a location with an address selected based on the outgoing stack pointer value. The valid stack capping value specifies a paging address indicating the paging of the address space including the address selected based on the outgoing stack pointer value in a predetermined portion of the valid stack capping value. The outgoing stack capping operation is complementary to the stack pointer switch validity check operation, where it sets a valid stack capping value on the outgoing stack, and the valid stack capping value is expected to pass the stack pointer switch validity check operation when the stack is switched back as the incoming stack in a later stack switch operation.

[0027] The outgoing stack capping operation does not necessarily depend on the incoming stack pointer meeting at least one stack capping value validity condition in the stack pointer switch validity check operation. Even if the incoming stack pointer is incorrect, resulting in the incoming data value loaded based on the incoming stack pointer value failing the check in the stack pointer switch validity check operation, it can still be safe to cap the outgoing stack with a valid stack capping value. In any of the cases described below, in some embodiments, the outgoing stack capping operation can be triggered by a different instruction from the stack pointer switch validity check operation. Thus, the circuit logic implementing the function of the instruction for triggering the outgoing stack capping operation can be defined independently of any validity check and can simply perform the outgoing stack capping operation. Even if an error is detected in response to the instruction triggering the stack pointer switch validity check operation, this does not necessarily stop the execution of the instruction for triggering the outgoing stack capping operation (the timing of responding to the error identified in the stack pointer switch validity check operation can vary and may not require immediate signaling of the error, as further explained below).

[0028] The paging address specified in the valid stack capping value for the outgoing stack capping operation can be a virtual paging address indicating the paging of the virtual address space including the virtual address selected based on the outgoing stack pointer value.

[0029] The techniques described above can be used in any of the following usage scenarios: where a stack data structure is used by a processing circuit to maintain thread-specific data, and different software threads can be allocated different stack data structures, such that it may be necessary to switch the stack pointer from the outgoing stack pointer to the incoming stack pointer during a thread switch to change which stack is active. The techniques can be particularly useful in usage scenarios where the data on the stack is sensitive or otherwise requires certain security measures to control the use of data on the stack, such that there is a risk during a stack pointer switch that switching to any address in the memory of a stack that does not represent having undergone those security measures may produce an incorrect result that can compromise the software code being executed.

[0030] However, a specific scenario in which this technique can be useful is one in which the stack pointer is a guard control stack (GCS) pointer that is used to control access to a guard control stack (GCS) data structure for protecting return addresses returned from an exception or a function call. This GCS data structure can be used as a defense against return-oriented programming (ROP) attacks, which are a common form of attack that aims to cause an incorrect control flow by changing the return address used upon a function return or an exception return (generally by changing the return address when it is stored in memory during a nested sequence of function calls or exceptions). The protected GCS data structure can be established in a region of memory that has at least one defense measure that limits the ability to write data to the GCS data structure, providing some additional protection relative to a normal memory region, such that it is less likely for an attacker to cause an unexpected instruction written to the GCS data structure to change the return state saved on the GCS data structure. The protected return state information from the GCS data structure can be used to directly control an exception / function return, or compared with exception / function return information obtained from software from other sources (e.g., a separate software-managed stack structure in memory) to check whether the exception / function return information can be safely used before triggering the corresponding exception / function return. However, although some security measures can be used to prevent incorrect updates to the GCS data structure, those measures can be ineffective if an attacker can easily bypass those security measures by switching the stack pointer used to access the GCS data structure to an incoming stack pointer value that points to a memory region controlled by the attacker, such that an incorrect return state is then returned when the victim software subsequently attempts to access the GCS data structure. Additionally, some attacks can be based on switching the incoming stack pointer to the correct GCS data structure, but switching to an incorrect location on the GCS data structure (e.g., a location that does not represent the current "top" of the stack) can create a risk that a subsequent exception / function return will behave incorrectly because it uses a return state intended for a different exception / function return.

[0031] Accordingly, a stack pointer switch validity check operation can be particularly useful in architectures that support the use of the GCS data structure because it implements a sanity check of whether the incoming GCS data structure can be safely used, without requiring every stack pointer switch to be trapped at a higher privilege level. Additionally, a more efficient encoding of the stack capping value using a paged address instead of a full address is useful because it frees up a large amount of stack record encoding for other purposes. Since the GCS data structure can be used to protect data from being tampered with by malicious attackers, it can be useful to provide encoding space for additional types of stack records to allow the GCS data structure to be used to protect other types of information (not only function / exception return state information).

[0032] Several security measures can be supported in the instruction set architecture used by the processing circuitry to protect the GCS data structure from tampering. The apparatus can have a memory management circuit that determines whether to permit access to a target address based on memory attribute data associated with the target address, where the memory attribute data specifies whether the target memory address space region that includes the target address is a GCS region for storing the GCS data structure, and where write access to the GCS region is restricted to a dedicated class of GCS access instructions. Thus, the GCS data structure can be allocated in a dedicated type of memory region (regarded by the memory management circuit as a different type of memory than normal memory used for general data or program code). By restricting the class of instructions that can be written to this GCS region of memory, this reduces the attack surface that an attacker can exploit when attempting to modify the return state saved on the GCS data structure.

[0033] The memory management circuit can reject a non-GCS store operation to the target address triggered by a store instruction other than the dedicated class of GCS access instructions in response to determining that the target memory address space region is the GCS region. By restricting the ability to write to the GCS region for certain GCS access types of store instructions, other more general store instructions cannot tamper with the contents of the GCS data structure, thus providing a greater security assurance for the protected return state information stored in the GCS data structure. This reduces the attack surface that an attacker can exploit when attempting to mount a ROP attack.

[0034] Similarly, the memory management circuit can reject a GCS load / store operation to the target address triggered by one of the dedicated class of GCS access instructions in response to determining that the target memory address space region is not the GCS region. Thus, access can be rejected for a memory region not specified as a GCS data structure if the access is triggered by an instruction of the GCS access type. This prevents instructions of the GCS access type from being misused to access regions of memory not intended for storing the GCS data structure, and gives confidence that a GCS read will be to a memory region that could not have been modified by non-GCS access instructions, for defense against ROP attacks. The loading of an incoming data value for a stack pointer switch validity check operation can be regarded as a GCS load operation, such that it fails if the address is identified by the memory attribute data as corresponding to a memory region type other than the GCS region type.

[0035] In some examples, when the memory attribute data specifies that the target memory address space region is the GCS region, the memory management circuit may also reject instruction fetches or branches to the target address. Thus, data stored in the GCS region cannot be considered executable instructions (although it may be considered an address pointer indicating another location in memory where executable instructions are stored). Given this property of the GCS region, it can be extremely useful to set a valid stack capping value (which is expected on the incoming stack in a region allocated as the GCS region type when performing a stack switching operation) to specify the page address of the same page that contains the address of the location storing the stack capping value itself. This reduces the risk of incorrect control flow resulting from accidentally interpreting the stack capping value as an exception return address or a function return address when accessing GCS data structures in the GCS memory region. If an attacker compromises the control flow in the victim code such that a function return or an exception return is triggered when the stack capping value is at the top of the stack, this can be a risk. Even if the attacker does cause this incorrect access to the stack, resulting in the stack capping value being returned as a protected return address for an exception return or a function return, the corresponding return operation will result in an identification error (identified either directly at the stack access or return operation, or later when attempting to fetch an instruction from the address represented by the stack capping value that was misinterpreted as a return address), because when the stack capping value points to its own page and this page is expected to be allocated as the GCS region type, the branch or instruction fetch is not allowed.

[0036] The condition on the incoming data value that specifies the page address is not the only condition that the incoming data value is required to pass to satisfy the check in the stack pointer switching validity check operation. Other conditions may also be imposed.

[0037] For example, in some embodiments, the at least one stack capping value validity condition further includes the condition that the least significant portion of the incoming data value has a bit pattern that cannot be specified by the least significant portion of any valid instruction address. More specifically, the at least one stack capping value validity condition may include the condition that the least significant portion has a specific bit pattern indicating a valid stack capping value, where the bit pattern cannot be specified by the least significant portion of any valid instruction address.

[0038] For example, this particular bit pattern may have at least one bit with a bit value of 1. For some processor architectures, instruction encodings may have a specific number of bits, and there may be a requirement that any instruction be stored at an address aligned to an address boundary that is a multiple of the instruction size. For example, an instruction may have a 32-bit encoding and may thus need to be stored at an address aligned to a 4-byte address boundary (for architectures with different instruction sizes, the address boundary may be at a different granularity). If an attempt is made to execute an instruction fetch or branch to an unaligned address, an error may be signaled. Thus, by encoding the stack cap value such that its least significant bit has a non-zero bit pattern that cannot be specified by the least significant part of any valid instruction address (because if it were treated as an instruction address, the pattern would make the value an unaligned instruction address), this can provide another measure for detecting an erroneous attempt to perform an exception return or a function return using the stack cap value as a return address, thus improving security.

[0039] Thus, a "stack cap value" assigned to an inactive stack (expected to be verified when switching back to that stack) can be considered to "cap" (or "seal") that stack, in the sense that a stack cap value is placed at the top of the stack to prevent the stack from being used for control return operations until the cap has been removed.

[0040] The incoming stack pointer value can be defined independently of the outgoing stack pointer value, such that the switch from the outgoing stack pointer value to the incoming stack pointer value can represent a switch to a completely different stack structure, rather than just an increment or decrement of the current stack pointer that occurs when pushing an entry onto the stack or popping an entry from the stack.

[0041] The processing circuit can perform the stack pointer switch validity check operation in response to a first stack pointer switch instruction specifying an operand indicating the incoming stack pointer value. Thus, the incoming stack pointer value can be defined as an arbitrarily specified operand of the first stack pointer switch instruction (e.g., by specifying the incoming stack pointer value in a general-purpose register specified by the first stack pointer switch instruction).

[0042] The apparatus may have a stack pointer register for storing the stack pointer. The outgoing stack pointer value for a stack pointer switch can be the current value specified for the stack pointer register when the first stack pointer switch instruction is executed.

[0043] In response to the first stack pointer switch instruction, when the stack pointer switch validity check operation is successful (i.e., the incoming data value obtained based on the incoming stack pointer value meets at least one stack cap value validity condition and satisfies any other checks (depending on the specific implementation)), the processing circuit can update the stack pointer register from the outgoing stack pointer value to the incoming stack pointer value.

[0044] In some examples, the first stack pointer switching instruction may be the only mechanism available to software (typically user-level software) executing in the lowest privilege execution state of the processing circuitry for switching the stack pointer. Software executing in the lowest privilege execution state may not have direct write access to the stack pointer register (it is possible that the software may still be able to directly write to the stack pointer register when executing in a higher privilege execution state). By restricting the ability of user-level software to cause a stack pointer switch, this improves security. Since the software must use the first stack pointer switching instruction to cause the stack pointer switch, this forces the execution of a stack pointer switch validity check operation. For example, when the current privilege level has a privilege lower than a certain threshold privilege level, a system register instruction that the software attempts to execute to request an update to the stack pointer register may be rejected.

[0045] In some embodiments, the first stack pointer switching instruction may also trigger the outgoing stack capping operation mentioned above. Thus, in response to the first stack pointer switching instruction, the current value specified for the stack pointer register when the first stack pointer switching instruction is executed is used as the outgoing stack pointer value, and a valid stack capping value that specifies the (virtual) paging address corresponding to the outgoing stack pointer value is formed, and the valid stack capping value is written to the location at the address corresponding to the outgoing stack pointer value.

[0046] However, in practice, performing both the stack pointer switch validity check operation and the outgoing stack capping operation in the same instruction may require translating two different target addresses: the load address of the incoming data value to be loaded based on the incoming stack pointer value, and the store address of the valid stack capping value to be stored based on the outgoing stack pointer value. Some circuit embodiments may not be able to handle two different address translations in response to the same instruction, and thus an instruction set architecture that requires performing both the stack pointer switch validity check operation and the outgoing stack capping operation in one instruction (although possible) can pose significant design challenges to microarchitecture circuit designers.

[0047] Therefore, to simplify circuit design embodiments, it may be useful to define two separate architectural instructions for performing stack switching: a first stack pointer switching instruction that triggers the stack pointer switch validity check operation and, if the validity check operation is successful, updates the stack pointer register to switch the stack pointer to the incoming stack pointer value; and a second stack pointer switching instruction that is used to perform the outgoing stack capping operation.

[0048] However, with the two-instruction method, the outgoing stack pointer value (replaced in the stack pointer register in response to the first stack pointer switch instruction) should be "remembered" between the first stack pointer switch instruction and the second stack pointer switch instruction. From an architectural perspective, it may not be desirable to consume architectural registers to preserve the outgoing stack pointer value as this can increase register pressure, and if architectural registers are allocated to store the outgoing stack pointer value to be used as an operand for the second stack pointer switch instruction, it may also not be desirable to expose the outgoing stack pointer value to software that executes after the stack pointer is switched, which can be a risk.

[0049] Accordingly, the instruction set architecture can define that the first stack pointer switch instruction architecturally writes the outgoing stack pointer value to the incoming stack (temporarily), such that the outgoing stack pointer value is preserved by writing it to a memory address space. This does not preclude some microarchitectural implementations that are able to use microarchitectural mechanisms (such as register renaming or store buffering) to provide the second stack pointer switch instruction with faster access to the outgoing stack pointer value nominally written to memory by the first stack pointer switch instruction, which is faster than what the second stack pointer switch instruction accessing memory would actually have (e.g., a buffer within the processing circuitry hardware can be used as it is expected that the second instruction will frequently need the outgoing stack pointer relatively soon after it is stored by the first instruction). Nevertheless, such microarchitectural circuit mechanisms are optional and from an architectural (software visible) perspective, these instructions can be processed as if the first stack pointer switch instruction writes the outgoing stack pointer value to memory and the second stack pointer switch instruction reads it from memory.

[0050] Accordingly, in response to the first stack pointer switch instruction, the processing circuitry can push an in-progress token value specifying the outgoing stack pointer value to a location having an address selected based on the incoming stack pointer value. This temporarily makes the outgoing stack pointer value accessible in the memory address space that is architecturally visible, such that it is acceptable to overwrite the outgoing stack pointer value in the stack pointer register to switch the stack pointer to the incoming stack pointer value, even if the outgoing stack has not been capped.

[0051] For the load operation to load the incoming data value to be verified in the stack pointer switch validity check operation, and the store operation to push the in-progress token value to the location at the address selected based on the incoming stack pointer value, the processing circuit can perform the load operation and the store operation atomically. Thus, a provision is made to prevent intervening accesses to the address selected based on the incoming stack pointer value between the load and store being executed, or if such intervening accesses are possible, then detect them and correct for this lack of atomicity (e.g., by triggering a pipeline flush and re-executing the first stack pointer switch instruction). Any known atomic access mechanism can be used to enforce atomicity between the load and the store.

[0052] In response to a second stack pointer switch instruction, the processing circuit can verify whether a given data value obtained by the memory access circuit in response to a memory access request specifying an address determined based on the current stack pointer in the stack pointer register is a validly formed in-progress token value; and in response to verifying that the given data value is a validly formed in-progress token value: trigger a write of a valid stack cap value to the location at the address determined based on a given stack pointer value specified by a portion of the given data value, the valid stack cap value specifying in a predetermined portion of the valid stack cap value a paging address indicating paging of an address space including the address determined based on the given stack pointer value. In some examples, the address determined based on the given stack pointer value can be an address offset from the given stack pointer value by an amount corresponding to the size of one stack entry. If the second stack pointer switch instruction is executed after the first stack pointer switch instruction, then the given data value expected to be obtained based on the current stack pointer (at the second stack pointer switch instruction) should be the in-progress token value previously written to the incoming stack data structure by the first stack pointer switch instruction. However, it should be understood that the hardware circuit logic of the processing circuit cannot know whether the second stack pointer switch instruction actually follows the first stack pointer switch instruction, i.e., the hardware circuit logic just implements "black box" circuit logic that implements the function of the second stack pointer switch instruction based on its defined operands, regardless of which other instructions are selected by software to execute before this instruction. Thus, there is no circuit in the hardware that will enforce that the operands of the second stack pointer switch instruction actually correspond to the operands used for the first stack pointer switch instruction.

[0053] Accordingly, the second stack pointer switching instruction checks that a given data value obtained based on the current stack pointer in the stack pointer register corresponds to a validly formed in-progress token value, and if valid, uses the given stack pointer value provided by the in-progress token value (expected but not guaranteed to be the address of the outgoing stack structure previously switched away from by a previous first stack pointer switching instruction) to access the corresponding stack structure, and pushes the valid stack cap value onto that structure. Similarly, a valid stack cap value is formed such that it specifies in a predetermined portion a paging address indicating paging of the address space including the address indicated by the given stack pointer value. This ensures that when the software later wishes to switch the corresponding stack back as the incoming stack for a stack pointer switching operation, a later check performed for the stack pointer switching validity check operation will pass.

[0054] In response to the second stack pointer switching instruction, when the given data value is verified to be a validly formed in-progress token value, the processing circuit is configured to update the current stack pointer in the stack pointer register to indicate removal of the entry that will provide the given data value from the corresponding stack data structure. For example, the update used to indicate entry removal may be the same type of update to the current stack pointer that would be performed when popping an entry from the stack (e.g., increment or decrement, depending on whether the stack grows / shrinks on push / pop). This update means that the in-progress token value will no longer be accessed on subsequent pop accesses to the stack. It is not necessary to actually delete the in-progress token value from memory because the update to the stack pointer means it will not be accessed on subsequent stack pointer-based accesses. If there is a subsequent push operation to push a new entry onto the stack, the in-progress token value may be overwritten later in any case.

[0055] In response to the second stack pointer switching instruction, when the least significant portion of the given data value has a bit pattern that cannot be specified by the least significant portion of any valid instruction address and cannot be specified by any value that meets at least one stack cap value validity condition, the processing circuit may verify that the given data value is a validly formed in-progress token value. Similarly, it may be useful for the in-progress token value to have a least significant portion corresponding to an unaligned instruction address such that any attempt to use the in-progress token value as a return address will result in an unalignment error as discussed above. Additionally, it may be necessary to have different encodings for the least significant portion of the valid stack cap value and the valid in-progress token value to avoid an attacker bypassing protection using miscontrolled flow (e.g., this allows detection of the error of only executing one of the first / second stack pointer switching instructions instead of both together).

[0056] In response to the second stack pointer switching instruction, the processing circuit may trigger an error handling response in response to determining that the given data value is not a validly formed in-progress token value.

[0057] For both the error handling response when an incoming data value fails at least one stack capping value validity condition of the stack pointer switching validity check operation, and the error handling response in response to a second stack pointer switching instruction when a given data value is not a validly formed in-progress token, the error handling response can be implemented in several ways. For example, the error handling response can include at least one of the following: signaling an error; setting an error reporting indicator; and setting the stack pointer to an invalid value. In some embodiments, the error handling response can directly cause the processing to be paused (e.g., signaling an error or exception upon detection of an error). Other examples can use more indirect means of notifying of the error such that the error itself may not occur until later. For example, by setting the stack pointer to an invalid value that cannot be a valid memory address, subsequent attempts to access memory based on the invalid stack pointer may trigger an error. This can sometimes be more easily implemented in hardware than directly triggering an error based on the checks of the first / second stack pointer switching instructions. For example, a reserved portion of the address space where the higher address bits have a value other than all 0s or all 1s can be reserved such that it cannot be specified as a valid address, so if an error handling response is needed, then by setting the stack pointer to an address in this reserved portion, this can cause a later memory access error such that the error handling can still be avoided. It should be understood that the specific error handling response taken can vary between different embodiments of the general techniques described above. Additionally, in some cases, the error handling response taken for a failed stack capping value validity check can be different from the error handling response taken for a second stack pointer switching instruction when a given data value is not a validly formed in-progress token.

[0058] The techniques discussed above can be implemented within a data processing apparatus having hardware circuitry provided for implementing the memory access circuitry and processing circuitry discussed above.

[0059] However, the same technique can also be implemented within a computer program that executes on a host data processing device to provide an instruction execution environment for the execution of object code. Even if the host data processing device itself does not support the architecture, this computer program can control the host data processing device to simulate the architectural environment that it would provide on a hardware device that actually supports object code according to a certain instruction set architecture. The computer program can have memory access program logic and processing program logic that control the host data processing device to emulate the features discussed above, including support for the stack pointer switch validity check operation. The memory access program logic performs stack access based on the stack pointer. The processing program logic performs the stack pointer switch validity check operation, which includes a check of whether a predetermined portion of the incoming data value loaded based on the incoming stack pointer value specifies a given paging address that includes a paging of an address derived from the incoming stack pointer value (in the emulated address space, rather than the host address space of the host data processing device). Thus, when object code that requires a stack pointer switch from an outgoing stack pointer value to an incoming stack pointer value is executed in the instruction execution environment provided by an emulated computer program executing on a host data processing device, the same functionality as discussed above can be achieved even if the host data processing device itself does not support the stack pointer switch validity check operation in hardware.

[0060] For example, such an emulator program can be useful when program code written for one instruction set architecture is executed on a host processor that supports a different instruction set architecture. Furthermore, since the execution of software on an emulated execution environment can enable software testing to proceed in parallel with the ongoing development of hardware devices that support a new architecture, emulation can allow software development for a newer version of an instruction set architecture to begin before the processing hardware that supports that new architecture version is ready. The emulator program can be stored on a storage medium, which can be a non-transitory storage medium.

[0061] Figure 1An example of a data processing apparatus 2 is shown schematically. The data processing apparatus has a processing pipeline 4 comprising a number of pipeline stages. In this example, the pipeline stages include a fetch stage 6 for fetching instructions from an instruction cache 8; a decode stage 10 for decoding the fetched program instructions to produce micro-operations (decoded instructions) to be processed by the remaining stages of the pipeline; an issue stage 12 for checking whether the operands required for the micro-operations are available in a register file 14 and for issuing that a micro-operation for performing the required operands of a given micro-operation once is available; an execution stage 16 for performing data processing operations corresponding to the micro-operations by processing the operands read from the register file 14 to produce result values; and a write-back stage 18 for writing the processed results back to the register file 14. It should be understood that this is only one example of a possible pipeline architecture and that other systems may have additional stages or different stage configurations. For example, in an out-of-order processor, a register renaming stage may be included for mapping architectural registers specified by program instructions or micro-operations to physical register specifiers that identify physical registers in the register file 14. In some examples, there may be a one-to-one relationship between the program instructions decoded by the decode stage 10 and the corresponding micro-operations processed by the execution stage. There may also be a one-to-many or many-to-one relationship between the program instructions and the micro-operations such that, for example, a single program instruction may be split into two or more micro-operations, or two or more program instructions may be fused to be processed as a single micro-operation.

[0062] Execution level 16 (an example of a processing circuit) includes a number of processing units for performing different categories of processing operations. For example, the execution units may include a scalar arithmetic / logic unit (ALU) 20 for performing arithmetic or logical operations on scalar operands read from register 14; a floating-point unit 22 for performing operations on floating-point values; a branch unit 24 for evaluating the result of a branch operation and adjusting the program counter, which represents the current execution point accordingly; and a load / store unit 26 for performing load / store operations to access data in the memory systems 8, 30, 32, 34. The load / store unit is an example of a memory access circuit. A memory management unit (MMU) 28 is provided, which is an example of a memory management circuit, for performing address translation between a virtual address specified by the operand of a data access instruction by the load / store unit 26 and a physical address identifying the storage location of the data in the memory system. The MMU has a translation lookaside buffer (TLB) 29 for caching address translation data from the page tables stored in the memory system, where the page table entries of the page tables define the address translation mapping and access permissions, which, for example, govern whether a given program being executed on the pipeline is allowed to read from or write to data or execute instructions from a given memory region. When traversing the page table structure to locate the page table entry corresponding to the required address, the MMU 28 may have circuitry for requesting memory access during the page table walk. The memory management unit is an example of a memory management circuit.

[0063] In this example, the memory system includes a level-1 data cache 30, a level-1 instruction cache 8, a shared level-2 cache 32, and a main system memory 34. It will be understood that this is only one example of a possible memory hierarchy, and other arrangements of caches may be provided. The specific types of processing units 20 to 26 shown in execution level 16 are only one example, and other specific implementations may have different sets of processing units or may include multiple examples of the same type of processing unit, such that multiple micro-operations of the same type can be processed in parallel. It should be understood Figure 1 is only a simplified representation of some components of a possible processor pipeline implementation, and a processor may include many other elements not shown for the sake of brevity. Although Figure 1 a single processor core with access to memory 34 is shown, the apparatus 2 may also have one or more additional processor cores that share access to memory 34, where each core has a corresponding cache 8, 30, 32.

[0064] Figure 2Shows an example of a call to a function (labeled fn1 for ease of reference) and a return from the function. A function (also called a procedure) is a sequence of instructions that can be called from another part of the program, and when it is completed, it returns the process to the part of the program flow where the function was called. The same function can be called from multiple different locations in the program, and thus when the function is called, the function return address is stored so that the function return can distinguish to which address the program flow should return.

[0065] For example, as Figure 2 shown, a branch with the link instruction BLR can be executed at the point (represented by the address #add1) where the function is to be called, causing the program flow to branch to the instruction at the branch target address #add2 specified by the operand of the branch with the link instruction. The branch with the link instruction also causes the processing circuit to set the link register (a specified register for tracking the function return address) to the address of the next instruction after the branch with the link instruction (in this example, the function return address is #add1 + 4). After the branch has been taken, several instructions (such as LD, MUL, ADD, etc.) are executed in the function code, and when the function is completed, a return branch instruction RET is executed, which branches to the instruction indicated by the return address stored in the link register.

[0066] If no other function is called from within fn1 and no exception occurs before reaching the return branch at the end of fn1, then the address in the link register should still be the same as it was set when fn1 was called.

[0067] However, typically, the first function fn1 called by the background program code itself can call another function (i.e., fn2) in a nested manner, and in this case, the function call to fn2 will overwrite the return address stored in the link register. Therefore, before calling another function, the function code of the first function fn1 should contain instructions to store the return address from the link register into a data structure in memory (such as a stack structure, operating in a last-in, first-out (LIFO) manner), and after returning from fn2, the function code of fn1 should restore the return address to the link register before the return branch is executed. The responsibility for saving and restoring the function return state (such as the return address) will typically lie with the software (there may be no architecturally enforced hardware mechanism for saving the return address).

[0068] However, when the function return address is stored in memory, it may be vulnerable to an attacker modifying the data, such as using another thread executing on another processor core, or by interrupting the calling function and simultaneously executing other program code that overwrites the return address stored in memory. Alternatively, the attacker can execute some instructions that target modifying the address operand of the instruction that restores the return address from memory to a register, such that the data loaded from memory is different from the return address that was originally stored to memory before the nested function call. If the attacker can cause the return branch to branch to a point in the program flow that is not an instruction after the function call branch, then the attacker may be able to cause the software to malfunction and may be able to bypass certain security protections or cause the execution of undesired operations.

[0069] Function calls are examples of operations that produce return status information that provides information about the state to which the processing circuitry will later be restored. Another context in which return status information can be captured is when an exception is taken, where an exception handling circuit in hardware, or a software exception handler, can capture the exception return status information, such as an exception return address that indicates the address of the instruction that will be executed after returning from handling the exception, and / or processor state information that stores the mode or execution state that the processor will be in after returning from the exception. For example, the stored processor state information can indicate the exception level at which the exception was taken, as well as other information about the operating state of the processor when the exception was taken. Like function calls, exceptions can be nested, and thus when another exception is taken, the exception return status captured for the exception can be stored to memory (automatically by hardware, or by a software exception handler), and thus may be vulnerable to attacker tampering when stored in memory. These types of attacks can be referred to as return-oriented programming (ROP) attacks. It may be desirable to provide architectural countermeasures against such attacks.

[0070] Figure 3 A method for protecting against ROP attacks using a protected data structure 40 called a "Guard Control Stack (GCS)" in memory is shown. The location of the GCS data structure within the memory address space can be selected by software, but the hardware provides architectural features designed to protect the GCS data structure from being tampered with by malicious attackers.

[0071] As Figure 1As shown, register 14 may include a control register that includes one or more protected control stack pointer (GCSPR) registers 36 for storing stack pointers, which indicate addresses on the GCS data structure. In some examples, the GCS pointer registers may be a banked register set that provides, respectively, for at least two execution states (e.g., exception levels) to implement different GCS structures within a software-reference memory operating in different execution states without requiring reprogramming of a shared stack pointer register after each transition of the execution state. Other examples may use a single GCS pointer register, and software may update the stack pointer stored in the GCS pointer register upon a transition between execution states.

[0072] As Figure 3 As shown, the GCS data structure 40 is stored in a memory region of a GCS area of memory designated by memory attributes directly or indirectly specified by associated page table entries of a page table used by a memory management unit (MMU) 28 for controlling address translation and access rights checking. The GCS area attributes may be specified directly within the encoding of the corresponding page table entry of a memory region that includes at least a portion of the GCS data structure, or may be indirectly referenced within a register referenced by the page table entry.

[0073] When a memory region is identified as a GCS region, then when a particular subset of GCS access instructions is executed, write access to that region is limited to write requests triggered by processing circuitry 16. General memory instructions used by software for general storage operations not intended to access GCS structures are not considered to be one of the restricted subset of GCS access instructions. The MMU 28 may still allow reading of GCS structures using general load instructions that result in the issuance of read requests that are not GCS memory access requests. When a memory access request requests access to a GCS region, the request is a write request and the request is not a GCS memory access request triggered by one of the restricted subset of GCS access instructions, then the memory access request is rejected and an error is signaled. The subset of GCS access instructions may include at least one GCS push instruction that causes return status information (e.g., a function return address from a link register, or an exception return address or a stored processor state captured when an exception is taken) to be pushed onto a location on the GCS structure determined by a stack pointer indicated by the GCS pointer register 58. The GCS push instruction also causes the stack pointer to advance by an amount depending on the size of the stack frame pushed onto the GCS (e.g., if the GCS is managed as an ascending stack, then increment the stack pointer by the size of the stack frame; or if the GCS is managed as a descending stack, then decrement the stack pointer by the size of the stack frame). The GCS access instructions may also include at least one form of GCS pop instruction that pops protected return information from the GCS structure. As well as a return instruction that returns the popped value from the stack, the GCS pop instruction also causes the stack pointer to be adjusted in the opposite direction to the direction in which the stack pointer was adjusted for the GCS push instruction (e.g., if the GCS is managed as an ascending stack, then decrement the stack pointer by the size of the stack frame; or if the GCS is managed as a descending stack, then increment the stack pointer by the size of the stack frame). As described below, the GCS access instructions may also include GCS stack pointer switch instructions GCSSS1, GCSSS2 that may allow special purpose values to be written to the stack in the GCS memory region for the purpose of protecting the stack from inappropriate switching of the stack pointer in GCSPR 36.

[0074] GCS access instructions may not be permitted to access memory regions not specified as GCS region types by the page table attributes. Thus, when the target memory region of such an access is not marked as a GCS region type, if an attempt is made to perform a GCS access (including memory accesses for the GCS stack pointer switching instructions GCSSS1, GCSSS2), an error may be signaled. By prohibiting the use of GCS access instructions to access non-GCS regions, this discourages programmers from using GCS access instructions unless they truly intend a GCS access, in order to reduce the attack surface available to an attacker. Additionally, this gives confidence that data accessed by a GCS pop instruction or verified by one of the GCS stack pointer switching instructions GCSSS1, GCSSS2 cannot be modified by non-GCS instructions.

[0075] The GCS structure is separate from any data structures used by software to maintain stored return status information in memory to handle nesting of function calls or exceptions. Thus, when function calls or exceptions are nested, the GCS structure is not intended to eliminate the need for software itself to keep track of the saving and restoring of return status information (the software-triggered saving of the return status can continue in the same manner as on a processor that does not support the GCS protection architecture measures discussed above). Instead, the GCS structure provides a region of protected memory that is protected from corruption by malicious program code and that can be used to provide information for verifying return status information intended to be used by software to return from a function call or exception handling.

[0076] In some embodiments, a GCS pop instruction that pops protected return status information from the GCS structure may also cause the processing circuitry 16 to compare the popped return status with the current return status information stored in a register (e.g., the link register for function return, or the exception return address register and / or the stored processor status register for exception return), and if there is a mismatch between the return status information popped from the GCS structure 40 and the expected return status information that the software intends to use for function / exception return, an error is signaled. Thus, software can be protected from corruption by including examples of GCS push and GCS pop instructions within the program code executed when a function call / return or exception entry / return occurs.

[0077] Other embodiments may define a separate instruction for verifying whether the desired return status information is valid, different from the instruction that pops return status information from the GCS structure 40.

[0078] Alternatively, the GCS pop instruction can directly pop the protected return status from the GCS into one or more registers that specify the return status for an exception return or a function return (or can be combined with an exception / function return instruction to pop the protected return status and use that status to control the exception / function return). In this case, there is no need to verify whether the expected return status information provided by the software is valid, because in this particular implementation, the GCS-protected return status is directly used to control the exception / function return. For example, for the GCS protection of the function return address, the function return address can be directly popped into the link register, replacing any software-managed function return address that the software may place therein based on its own managed stack structure.

[0079] In addition, other types of GCSs for access instructions can also be supported. When the GCS mode is enabled (the control status in the control register can control whether the GCS mode is enabled), some instructions that have other functions in a mode where the use of the GCS is disabled can cause the processing circuit 16 to perform additional functions (such as additional GCS mode-specific security checks) when executed.

[0080] Generally, by providing architectural support for defining the types of GCS memory regions for the GCS structure 40 and restricting write access to the GCS region types to a limited subset of GCS access instructions that may not be allowed to access memory regions other than the GCS region types, this reduces the attack surface that an attacker can use to attempt to tamper with the protected return status information stored on the GCS structure 40.

[0081] Figure 4It is a flowchart showing the access permission check for GCS load / store operations. The GCS load / store operations are load / store operations triggered by one of the categories of GCS access instructions (e.g., GCS access instructions include GCS push and pop instructions, and GCS stack pointer switching instructions GCSSS1, GCSSS2). In step 110, the processing circuit 16 determines the target address of the GCS load / store operation based on the GCS pointer in register 36. In step 112, the memory management unit 28 looks up the target address in its TLB 29 to obtain the memory attributes of the target address of the GCS load / store operation. In step 114, the MMU 28 determines whether the target address corresponds to a GCS memory area, which is a dedicated type of memory area for storing GCS data structures. If the target address does not correspond to the GCS memory area type, then in step 116, the GCS load / store operation is rejected. An error is signaled, which may interrupt the currently executing process and cause the exception handler to handle the cause of the error. By suppressing GCS accesses to areas not marked as GCS memory area types, this prevents GCS load / store instructions from being misused to access non-GCS memory, and also means that the protected return state returned by a GCS load operation can be trusted because it could not have been tampered with by non-GCS instructions.

[0082] If in step 114, it is determined that the target address corresponds to a GCS memory area, then in step 118, the MMU 28 determines whether any other access permission checks pass. These checks may examine other attributes, such as other attributes indicating whether read requests and write requests are allowed respectively as the read / write permission information of the memory area, or attributes defining a subset of the execution state of the processor 2 where access to the area is allowed. If any other access permission check fails, then in step 116, the GCS load / store operation is rejected and an error is signaled. The error type information set by the processor when an error occurs may vary depending on whether the cause of the error is a GCS access to a non-GCS memory area or another type of access permission violation. If all any other access permission checks pass, then in step 120, the GCS load / store operation is allowed.

[0083] Figure 5 It shows a similar access permission check performed for non-GCS load / store operations (load / store triggered by instructions other than the GCS access instruction category). Steps 130, 132, and 134 are similar to Figure 4 steps 110, 112, 114. Compared with Figure 4 , in Figure 5In step 134, compared with GCS load / store operations, the response of non-GCS load / store operations to the check of whether the target address corresponds to the GCS memory area is in the opposite direction. When the target address corresponds to the GCS memory area, then the non-GCS store operation is rejected. If the target address does not correspond to the GCS memory area, then the GCS store operation is rejected.

[0084] Therefore, if in step 134 it is determined that the target address corresponds to the GCS memory area, and in step 135 it is determined that the current load / store operation is a non-GCS store operation, then in step 136, the non-GCS load / store operation is rejected and an error is signaled. Depending on the result of any other access permission check performed in step 138, even if the target of the non-GCS load operation is the GCS memory area, these non-GCS load operations may still be allowed. If any other access permission check fails, then in step 136, the non-GCS load / store operation is rejected again. Otherwise, if the target address does not correspond to the GCS memory area (No in step 134), or the non-GCS operation is a load operation (No in step 135), and in step 138 any other access permission check (unrelated to the GCS access check) passes, then in step 140, the non-GCS load / store operation is allowed.

[0085] Figure 6 is a flowchart showing the permission check for instruction fetch or branch operations. In step 150, the fetch stage 6 requests to fetch an instruction associated with the target address, or detects a branch to the instruction at the target address (which can be detected based on executing a branch instruction through the branch unit 24 of the execution stage 16, or based on a future branch prediction made by the branch predictor associated with the fetch stage 6). In response to the instruction fetch or branch, in step 152, the MMU 28 looks up the target address in its TLB 29 to identify the memory attribute data corresponding to the target address (where if there is a TLB miss, then a page table walk to memory is performed), and determines whether the target address is in the area identified as the GCS area by the memory attribute data. If the target address is in the GCS area, then in step 154, the instruction fetch or branch is rejected and an error is signaled. This prevents data on the GCS structure from being treated as executable instructions that may pose a risk of unpredictable results.

[0086] If the target address is not located within the GCS region, then in principle the instruction loaded from that address can be executed, depending on any other privilege checks performed at step 156. For example, these checks can include page table-based checks that examine the execution privilege data indicated in the memory attribute data at the target address. If those other privilege checks pass, then at step 158, instruction fetch or branching can be permitted. If the other privilege checks fail, then at step 154, instruction fetch or branching is again denied and an error is signaled.

[0087] Accordingly, the GCS memory region is limited to providing data values (such as address pointers and other information) and cannot provide executable instruction codes, because attempting to fetch an instruction from it, or branch an instruction to it, will trigger an error in the GCS region.

[0088] The measures described above can be useful for protecting the content on the GCS data structure from being tampered with by malicious attackers when processing function returns and exception returns within a particular program.

[0089] However, another possible risk to the control flow can be when the GCS pointer in register 36 switches from an outgoing stack pointer value associated with an outgoing stack data structure to an incoming stack pointer value associated with an incoming stack data structure. Such a switch can be common (and permitted) when switching between threads that use different GCS data structures, but can be a route that an attacker might attempt to use to cause improper processing when attempting to expose sensitive information that should not be accessible to the attacker but is accessible to the victim program. For example, if the incoming stack pointer value for the stack pointer switch is caused to be an address that does not correspond to the correct GCS data structure for that incoming thread, then there can be a risk that an improper control flow is triggered, especially if the address specified for the incoming stack pointer value happens to be an address within the GCS region of memory (e.g., an address that points to a return status on a GCS structure associated with a different thread). Even if the incoming stack pointer value is an address within the correct GCS data structure for that incoming thread, if the incoming stack pointer value is set to a location other than the location that represents the current "top" of the stack, then there can still be a risk of an error, because this can result in a valid return status for an exception / function return that is being used as a return address for a different exception / function return, which still results in incorrect information.

[0090] Such errors in the incoming stack pointer value can occur accidentally due to programming errors, or maliciously by an attacker who compromises the execution code to cause the value of the GCS pointer in register 36 to switch to point to an incorrect location (e.g., a location within a GCS data structure controlled by the attacker) (within the GCS region and thus not being Figure 4The instruction caught by the check is executed, and the incorrect location has been pre-filled with a return address designed to deceive the victim software program into branching to an incorrect code sequence.

[0091] One way to mitigate against such misconfigurations of the incoming stack pointer value could be to require that every write to GCSPR36 be trapped to a higher-privilege execution state, so that a higher-privilege program (such as an operating system or hypervisor) can verify whether the new value of the stack pointer is safe before updating GCSPR 36. However, this would adversely affect performance. Switching of the GCS pointer is very common when switching between application-level threads and ensuring fast switching and thus high performance may require allowing such GCS pointer switches to occur without calling a higher-privilege program.

[0092] It may also not be desirable to allow direct access to the GCS pointer register 36 to reduce the risk of attacks of the form described above.

[0093] Accordingly, some instructions can be provided for controlling the switching of the stack pointer value in the GCS pointer register 36, which can implement some sanity checks for detecting inappropriate configurations of the incoming stack pointer value. These GCS pointer switch instructions are considered to be members of the GCS access instruction class, such that writes to the GCS region triggered by these instructions can pass through Figure 4 and Figure 5 the checks shown.

[0094] When switching between different threads, the software stack (tracked using a stack pointer separate from the GCS pointer) is switched by software, and the current protected control stack can also be switched. To ensure that software cannot switch to any arbitrary location on the incoming GCS, a stack capping value is added to the outgoing GCS when switching out of that GCS, and this capping value is verified when switching to that GCS upon an incoming stack switch on the stack pointer.

[0095] The capping value is designed to be distinguishable from any procedure return value, which means that when switching stacks, we can be confident that we are only switching to the top of the incoming GCS. For example, the lower address bits of the capping value can be set to a non-zero bit pattern that cannot be specified by any valid instruction address, since instruction alignment requirements may require all valid instruction addresses to be aligned to an address boundary at the granularity of the instruction length (e.g., for 32-bit instruction encoding, the alignment address boundary could be a 4-byte interval). This means that any valid procedure return value should have 0 for the lower bits, and thus the capping value can be distinguished by having at least one 1 in the lower bits.

[0096] The capping value also contains the address of the capping location that we check when switching to the GCS. This address provides:

[0097] 1) We are using the correct VA of the PA mapping of the GCS to return to the sanity check of the expected GCS by comparing the capped address with the value loaded from the cap.

[0098] 2) A security check to prevent the case where a cap is (incorrectly) used as part of a procedure return. If we attempt to branch to the value held in the cap, then this implicitly is a *data* address as it is in the GCS region (so if there is an attempt to execute an instruction at that address, then Figure 6 the check shown will trigger an error), and when the processor attempts to execute from this location, it will not have execution permission and will take an exception.

[0099] We observe that if the address held in the cap is in the same memory page as the location of the cap, then the security check in point 2 above continues to hold as the address will still be a *data* page (the GCS memory region which is set using memory attributes directly or indirectly specified by page table entries defined by the granularity of the paging of the memory address space). If we only store bits [63:12] of the location of the cap, then this allows us to store fewer address bits as this ensures we are in the same 4KB page as the cap location. Thus, attempting to use the cap as a return address will still result in a permission error.

[0100] Sanity check #1 is also partially retained where we are confident that we are switching to the same GCS as expected, although only to a location within the same GCS (but in any case, it is still not possible to effectively switch to a location in that GCS that is not at the top of the stack when the GCS was previously switched out as switching to a location that is not at the top of the stack would result in information other than the valid cap value being returned, thus causing an invalid stack cap value validity check and hence an execution error handling response).

[0101] Therefore, the cap value only needs to specify the page address and does not need to specify the sub - page address bits. When checking the validity of the cap value on switching to the incoming stack, the sub - page address bits of the incoming stack pointer do not need to be checked to detect if the cap value is valid.

[0102] This is useful as it allows the format of the records in the GCS to have more reserved bits, thus allowing for future augmentation of the architecture. This is very useful as the expected GCS can be used to protect a wide range of other information in addition to the function / exception return status information, and thus having encoding space for encoding other record types can be useful.

[0103] Figure 7Shows the steps performed when switching the stack pointer in the GCSPR 36 from an outgoing stack pointer value to an incoming stack pointer value (where the incoming stack pointer value is independent of the outgoing stack pointer value, i.e., this is not just an increment or decrement of the current stack pointer performed when pushing onto or popping from the stack, but rather any switch of the stack pointer to specify an address provided as an operand of an instruction triggering the switch).

[0104] In step 200, the processing circuit 16 performs a stack pointer switch validity check operation. This includes issuing a memory access request to load an incoming data value from an address X determined based on the incoming stack pointer value. In many cases, the address X may be the address indicated by the incoming stack pointer value itself. However, it will also be feasible that the address X of the incoming data value is determined by applying an offset to the address indicated by the incoming stack pointer value, depending on how the position marked as the "top" of the stack is represented relative to the stack pointer value in register 36. The memory access request to load the incoming data value from address X is a GCS load access, so it may fail if address X does not correspond to GCS memory.

[0105] If the software is operating correctly, the incoming data value expected to be obtained from address X based on the incoming stack pointer value should be the stack capping value that was previously placed on the incoming GCS structure when switching out of that GCS structure on the previous switch of the GCS pointer. However, the processing circuit 16 does not know (yet) whether the incoming stack pointer value actually corresponds to a previously accessed GCS structure, so it is also possible that the incoming data value may not correspond to a valid stack capping value.

[0106] Therefore, in step 200, the processing circuit 16 performs a stack pointer switch validity check operation to check whether the incoming data value meets at least one stack capping value validity condition. At least one stack capping value validity condition at least includes that a predetermined portion of the incoming data value corresponds to a given paging address indicating the paging of the address space including address X, regardless of whether another portion of the incoming data value corresponds to the sub-paging address bits of address X. More specifically, this predetermined portion of the incoming data value must correspond to the virtual paging address corresponding to the virtual address X, such that this check can verify whether the virtual-to-physical address mapping used to access the stack based on the incoming stack pointer value is still the same as when the stack was previously accessed. By not considering the sub-paging address bits of address X for verifying at least one stack capping value validity condition, this frees up a significant amount of encoding space in the GCS stack record for other purposes, as will be discussed in more detail below regarding Figure 11 discussed in more detail below Figure 8 As shown in step 224 below, at least one stack capping value validity condition may also impose other requirements on the encoding of the incoming data value.

[0107] In step 202, it is determined whether the incoming data value complies with each of at least a stack capping value validity condition imposed by the processing circuitry 16. The at least one stack capping value validity condition to be satisfied includes at least a virtual paging address where the incoming data value specifies a location for storing the incoming data value. One or more additional stack value validity conditions may also be imposed (e.g., the lower portion of the incoming data value is a token bit pattern representing a stack capping value). If any of the stack capping value validity conditions are not met, then in step 204, an error handling response is triggered. For example, this response may signal an error, set an error flag in an error reporting register, or set the stack pointer in the GCSPR 36 to an invalid address within a reserved address range that cannot be specified for any valid memory address (if there is any attempt to perform a load / store operation or instruction fetch on an address within that reserved range, the MMU may trigger an error).

[0108] If it is determined that the incoming data value does not comply with each of the at least one stack capping value validity condition imposed, then in step 206, the processing circuitry 16 allows the stack pointer to switch from the outgoing stack pointer value to the incoming stack pointer value (so the incoming stack pointer value is written to the GCSPR 36).

[0109] In addition, in step 208, the processing circuitry 16 performs an outgoing stack capping operation to push a valid stack capping value into a location at an address Y selected based on the outgoing stack pointer value. The valid stack capping value specifies, in a predetermined portion to be checked in step 200, a (virtual) paging address indicating paging of the address space including (virtual) address Y. Thus, this ensures that the outgoing stack is sealed (to prevent a return from being triggered based on access to the top of the stack, since the stack capping value has a value that causes an instruction fetch or branch to that address to result in an error), and has an appropriate capping value that will later allow a valid switch back to that stack when the corresponding stack pointer value is specified as the incoming stack pointer for stack switching.

[0110] In some embodiments, Figure 7 the outgoing stack capping operation shown may be performed in response to the same instruction that also triggers a stack pointer switch validity check operation.

[0111] However, Figure 8 and Figure 9 illustrates a specific example where the outgoing stack capping operation is performed in response to an instruction different from the instruction that triggers the stack pointer switch validity check operation. Separating these operations into separate instructions may significantly simplify the circuit implementation, as it avoids the need for two different memory addresses to be translated by the MMU 28 in the same instruction. Figure 8Shows the function of the first stack pointer switch instruction (GCSSS1) that performs a stack pointer switch validity check operation and implements the switch of the stack pointer in GCSPR36. Figure 9 Shows the function of the second stack pointer switch instruction (GCSSS2) that performs an outgoing stack capping operation.

[0112] As Figure 8 shown in step 220, the operands of the first stack pointer switch instruction are the outgoing stack pointer value specified in GCSPR 36 and the incoming stack pointer value specified in the general-purpose register Xn, and the register specifier of the general-purpose register Xn is encoded within the instruction encoding of the first stack pointer switch instruction. Thus, software can select any general-purpose register for defining the incoming stack pointer, and the previous instruction can set the incoming stack pointer in any arbitrary software-specific manner. GCSPR36 does not need to be explicitly encoded in the encoding of the GCSSS1 instruction because it is an implicit operand of the operation.

[0113] In response to the GCSSS1 instruction being decoded by decode stage 10 and issued for execution by issue stage 12 of pipeline 4, at step 222, processing circuit 16 controls the memory access circuit (load / store unit 26) to issue a load memory access request to load an incoming data value from a location corresponding to the address determined based on the incoming stack pointer value (derived from register Xn). In one particular embodiment, the address specified by the load memory access request is equal to the incoming stack pointer value, but other examples can derive the load address by applying an offset to the incoming stack pointer value. This load memory access request depends on Figure 4 the checks shown, since it is a GCS load operation as it is triggered by the GCSSS1 instruction. Assuming the load operation passes these checks, the incoming data value is returned, and at step 224, processing circuit 16 checks (i) whether a predetermined portion of the incoming data value corresponds to the paging address portion of the virtual address for the load in step 222, and (ii) whether the least significant portion of the incoming data value is a specific bit pattern that identifies a valid stack capping value, and this specific bit pattern cannot be specified by the least significant portion of any valid instruction address. If either of these checks fails, then the incoming data value is not a valid stack capping value, and at step 226, an error handling response is triggered (which can be any of the types of error handling responses mentioned above for Figure 7 step 204).

[0114] If the check in step 224 passes (a predetermined portion of the incoming data value corresponds to the paging address portion of the virtual address used for the load in step 222, and the least significant portion of the incoming data value has a specific bit pattern identifying a valid stack capping value), i.e., the incoming data value meets at least one stack capping value validity condition, then in step 228, the processing circuit 16 controls the memory access circuit 26 to issue a store memory access request to store the in-progress token value to a memory location having an address selected based on the incoming stack pointer. The in-progress token value specifies the outgoing stack pointer value and has its least significant portion set to another bit pattern that cannot be specified by the least significant portion of any valid instruction address, which is different from the encoding of the least significant portion of the valid stack capping value. It may not be expected that it is necessary to use the outgoing stack pointer as an operand for an instruction that switches the stack pointer because it would be expected to be feasible to simply overwrite the outgoing stack pointer with the incoming stack pointer value in the GCSPR 36. However, by writing the outgoing stack pointer to the incoming stack (after the incoming stack has been verified by performing the stack capping value validity check of step 224), this helps support the implementation where the stack pointer switch operation is split between two instructions, GCSSS1 and GCSSS2, because it enables the outgoing stack pointer value to be saved so that it can be used for the second stack pointer switch instruction, GCSSS2, without consuming an architectural register to retain the outgoing stack pointer value after it has been overwritten in the GCSPR 36.

[0115] In step 230, in the case where the check in step 224 passes, update the GCSPR 36 to specify the incoming stack pointer value (thus overwriting the outgoing stack pointer value).

[0116] As Figure 8As shown, the load at step 222 and the store at step 228 are performed atomically, such that the result seen by software executing the GCSSS1 instruction and another thread accessing the same address determined based on the incoming stack pointer is consistent with the result that would occur if the load at step 222 and the store at step 228 were executed sequentially without any intervening writes to that address occurring between the load at step 222 and the store at step 228. There may be several different techniques available to ensure this atomicity. For example, a lock-based method may be used to lock access to the relevant memory location such that no other writes to that memory location are allowed during the period between the load and the store. Alternatively, conflict write operations triggered by other threads (e.g., executing on a different processor core) may still be allowed during the intervening period, but if such a write is detected, a mechanism may be provided to detect such conflict write operations and abort the processing executed by the stack pointer switch instruction (e.g., clearing the pipeline and winding back to the previous execution point in the thread containing the GCSSS1 instruction, causing the GCSSS1 instruction to be re-executed later). The atomic processing of the load and the store helps improve security by reducing the chance that the state checked by the GCSSS1 instruction changes (which may be at risk of being incorrect) before that check is complete.

[0117] Figure 9 Illustrates the processing of a second stack pointer switch instruction (GCSSS2). As shown in step 250, the operand of the instruction is the current value specified in GCSPR 36. In the common usage scenario where the GCSSS2 instruction follows the GCSSS1 instruction, it is expected that the value in GCSPR36 should be the incoming stack pointer value specified based on the Xn operand of the GCSSS1 instruction. However, in practice, it is not possible for the hardware of the processor to check that the GCSSS2 instruction actually follows the GCSSS1 instruction. Thus, although the following pseudocode for GCSSS2 refers to the operand of the GCSSS2 instruction as the "incoming" stack pointer value, more generally, it operates on the current stack pointer stored in stack pointer register 36, regardless of whether this is actually the same value as the incoming stack pointer value of the previous GCSSS1 instruction.

[0118] In response to the GCSSS2 instruction being decoded by decode stage 10 and issued for execution by issue stage 12 of pipeline 4, at step 252, processing circuit 16 controls memory access circuit 26 to issue a load memory access request to request to load a given data value from a location corresponding to an address determined based on the current stack pointer obtained from GCSPR 36. In one particular implementation, the address specified by the load memory access request is equal to the current stack pointer value from GCSPR 36, but other examples may derive the load address by applying an offset to the current stack pointer value. This load memory access request depends onFigure 4 The inspection shown, since it is a GCS load operation, is triggered by the GCSSS2 instruction. Assuming the load operation passes these inspections, a given data value is returned, and at step 254, processing circuitry 16 checks whether the given data value is a validly formed in-progress token. For example, when the least significant bit portion of the given data value matches a specific bit pattern used to encode an in-progress token, processing circuitry 16 determines that the given data value is a validly formed in-progress token. This bit pattern is different from the bit pattern at the least significant portion of a valid stack cap value and thus cannot be specified by any value that meets at least one stack cap value validity condition. Additionally, this bit pattern cannot be specified by any valid instruction address (e.g., because it is non-zero in lower bit positions, which are less significant than the bits representing the significance of the instruction alignment boundary). If the given data value is not a validly formed in-progress token, then at step 256, an error handling response is triggered (again, this can be any of the previously mentioned error handling response types).

[0119] If the given data value is a validly formed in-progress token, then at step 258, processing circuitry 16 controls memory access circuitry 26 to issue a store memory access request to write the valid stack cap value to a location at address Y determined based on a given stack pointer value specified in a portion of the given data value loaded at step 252. In some examples, address Y may be an offset from the given stack pointer value by an amount corresponding to the size of one stack entry to reflect that adding the valid stack cap value to the top of the stack implies that the stack pointer for the corresponding stack will increment or decrement from its previous value (represented by the given stack pointer value), similar to any other stack push operation that pushes a new stack entry onto that stack. Thus, the store operation at step 258 is to a different address Y than the address from which the given data value was loaded at step 252. If the GCSSS2 instruction is not preceded by a previous GCSSS1 instruction, then the store at step 258 is to the outgoing stack structure based on the outgoing stack pointer of the GCSSS1 instruction, while the store at step 252 is from the incoming stack structure. The valid stack cap value specifies the (virtual) paging address of the paging of the address space including virtual address Y within a predetermined portion of the paging address check at step 224 for the GCSSS1 instruction, where the virtual address Y is specified by the given stack pointer value obtained from the given data value loaded at step 252. This sets the stack cap value on the outgoing stack such that when the stack pointer is later switched back to the outgoing stack (now used as the incoming stack), this later stack pointer switch can pass the check performed at step 224 for the GCSSS1 instruction. Figure 8 of the GCSSS1 instruction. Figure 8

[0120] Additionally, at Figure 9 ​In step 260, the processing circuit updates the stack pointer in GCSPR36 in response to the GCSSS2 instruction to indicate removal of the entry that will provide the given data value from the stack. This update can be an increment or a decrement of the current stack pointer by an amount corresponding to the size of one stack record. Whether to apply an increment or a decrement depends on whether the stack is managed as a decreasing stack or an ascending stack (both options are possible, i.e., for a decreasing stack the stack pointer is decremented on a push and incremented on a pop, and vice versa for an ascending stack).

[0121] Example pseudocode for representing the functionality of the GCSSS1 and GCSSS2 instructions is shown below. It should be understood that although this is shown as pseudocode in a hardware processor, it will be implemented using hardware circuit logic gates that implement the corresponding functionality.

[0122]

[0123]

[0124] Of course, the specific encodings for the valid stack cover value and the in - progress token can be different from those shown in the pseudocode.

[0125] Figure 10 An example of stack switching based on the operations discussed above is schematically shown. Figure 10 The top part shows the initial state when the outgoing GCS pointer P1 in GCSPR 36 currently points to the top of the outgoing GCS, and the incoming GCS (associated with pointer P2) has previously been capped by storing a valid stack cover value at the top of the stack.

[0126] As Figure 10 shown in the middle part, when the GCSSS1 instruction is executed to assign the operand indicating pointer P2 to the incoming GCS, the stack pointer switch validity check operation passes because the paging address part P2 of the incoming stack pointer value matches a predetermined part P2[63:12] of the value obtained from the position on the incoming GCS corresponding to the incoming stack pointer value. Thus, the incoming stack is marked with the in - progress token value that designates the outgoing stack pointer P1[63:3], and GCSPR 36 is updated to indicate the incoming stack pointer P2.

[0127] As Figure 10As shown in the lower part of, when the GCSSS2 instruction is executed, the in - progress token passed to the GCS (accessed based on the current stack pointer value P2 in GCSPR 36) is verified and passes the check. Thus, the outgoing stack is capped by writing the valid stack cap value to the location on the outgoing stack addressed based on the outgoing stack pointer P1[63:3] obtained from the in - progress cap value token. The valid stack cap value specifies the paging address part of the page including the location written with the valid stack cap value. Additionally, the current stack pointer in GCSPR 36 is updated to represent the removal of the in - progress token from the incoming GCS (in this example by adding 8, but other examples may use decrement or may have stack records of different sizes and thus different increment / decrement sizes).

[0128] Figure 11 Shows examples of different formats of GCS records that can be placed on the GCS stack:

[0129] ● Procedure (function) return address record 300, which provides a 62 - bit procedure return address.

[0130] The lower 2 bits of 0b00 indicate that this is a procedure return record (since these bits are 0, the address is 4 - byte aligned and thus can be a valid instruction address for a 32 - bit instruction encoded to be aligned to the address boundary).

[0131] ● Stack cap value record 302, which provides a 52 - bit paging address and has a value of 0b000000000001 for the lower 12 bits. Since the lower bits are 0b01, if interpreted as an instruction address, this would be an unaligned address and thus would trigger an alignment error.

[0132] ● In - progress token record 304, which provides a 61 - bit stack pointer value (to represent the outgoing stack pointer during stack pointer switching as mentioned above). With this encoding, there is a requirement for this outgoing stack pointer to be double - word aligned, such that the 62nd bit of the stack pointer address is implicitly 0. The lower 3 bits of record 304 have the encoding 0b101 to indicate that this is an in - progress token.

[0133] ● Exception record 306, which has its lower 5 bits set to 0b01001 and all more significant bits are 0. When the exception record 306 is detected on the stack, a defined number of subsequent records (where a full 64 bits can be used to record exception status information) provide items of the exception return status, e.g., the exception return address, the stored processor state associated with the main return of the process, etc.

[0134] This encoding means that there are several reserved encodings kept idle and available for future expansion:

[0135] ● Encoding 308 sets the lower 3 bits to 0b001 (the same as the stack capping record 302 and the exception record 306), but it sets the next 9 bits to anything other than 0b000000000 or 0b000000001. These encodings allow encoding at least 2 9 – 2 = 510 additional record types (by setting different “record type” indicators in bits [11:3]), each supporting the recording of other type information in the remaining 52 bits at bits [63:12] of the record. This additional encoding space is freed because the stack capping value record 302 specifies a paging address rather than a full address.

[0136] ● Encoding 310 sets the lower 2 bits [1:0] to 0b10 or 0b11, which supports a full 62-bit value specified in the remaining bits [63:2].

[0137] Figure 12 Shows a simulator implementation that can be used. While the earlier-described embodiments implement the present invention in terms of devices and methods for operating specific processing hardware that supports the technology of interest, it is also possible to provide an instruction execution environment in accordance with the embodiments described herein, which is implemented through the use of a computer program. Such computer programs are often referred to as simulators because they provide a software-based implementation of a hardware architecture. Types of simulator computer programs include emulators, virtual machines, models, and binary translators (including dynamic binary translators). In general, a simulator implementation can run on a host processor 1330 that optionally runs a host operating system 1320 and supports a simulator program 1310. In some arrangements, there can be multiple layers of simulation between the hardware and the provided instruction execution environment and / or between different instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide a simulator implementation that executes at a reasonable speed, but this approach can be justified in some cases, such as when it is necessary to execute program code native to another processor because of compatibility or reuse reasons. For example, a simulator implementation can provide an instruction execution environment with additional functionality not supported by the host processor hardware, or provide an instruction execution environment generally associated with a different hardware architecture. A review of simulation is given in “Some Efficient Architecture Simulation Techniques,” Robert Bedichek, Winter 1990 USENIX Conference, pages 53 to 63.

[0138] Where an implementation has been previously described with reference to specific hardware architectures or features, in a simulated implementation, equivalent functionality may be provided by a suitable software architecture or feature. For example, a specific circuit may be implemented as computer program logic in a simulated implementation. Similarly, memory hardware (such as registers or cache memory) may be implemented as software data structures in a simulated implementation. In arrangements where one or more of the hardware elements mentioned in the previously described implementations are present on host hardware (e.g., host processor 1330), some simulated implementations may utilize the host hardware as appropriate.

[0139] The simulator program 1310 may be stored on a computer-readable storage medium (which may be a non-transitory medium) and provide a program interface (instruction execution environment) to the target code 1300 (which may include application programs, operating systems, and hypervisors), the program interface being the same as the interface of the hardware architecture modeled by the simulator program 1310. Thus, program instructions of the target code 1300 that include stack pointer switching instructions GCSSS1, GCSSS2 as described above may be executed within the instruction execution environment using the simulator program 1310, such that the host computer 1330, which does not actually have the hardware features of the apparatus 2 discussed above, can emulate these features. Similarly, the various memory management check functions and triggers for accesses to memory (including support for GCS memory region types) as discussed above for the MMU 28 and load / store unit 26 may be emulated using the memory access program logic 1318 of the simulator program 1310.

[0140] Accordingly, the simulator program 1310 may have processing program logic 1312 that simulates the state of the processing circuit 4 described above. For example, the processing program logic 1312 may simulate a transition of the execution state in response to an event occurring during the simulated execution of the target code 1300 and perform processing operations. Instruction decoding program logic 1314 (which may be regarded as part of the processing program logic) decodes the instructions of the target code 1300 and maps these instructions to corresponding instruction sets in the native instruction set of the host device 1330. Register emulation program logic 1316 maps the register accesses requested by the target code to accesses to corresponding data structures maintained on the host hardware of the host device 1330, for example, by accessing data in the registers or memory 1332 of the host device 1330 (e.g., when an instruction (such as a GCS push / pop instruction or GCSSS1, GCSSS2 instruction) requires access to the GCS pointer register for execution, the access to the host location representing the simulated GCS pointer register can be managed by the register emulation program logic 1316). Memory access program logic 1318 performs address translation, page table traversal, and access control checks in a manner corresponding to the MMU 28 described in the hardware implementation above, but also has the additional function of mapping the simulated physical addresses (obtained through address translation based on the page table defined for the target code 1300) to host virtual addresses used to access the host memory 1332. These host virtual addresses can themselves be converted to host physical addresses using the standard address translation mechanism supported by the host (converting the host virtual addresses to host physical addresses is outside the scope controlled by the simulator program 1310).

[0141] In this application, the term "configured to..." is used to mean that an element of a device has a configuration capable of performing the defined operation. In this context, "configuration" means an arrangement or manner of interconnection of hardware or software. For example, the device may have dedicated hardware that provides the defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not mean that the device element needs to be changed in any way to provide the defined operation.

[0142] In this application, a list of features prefaced by the phrase "at least one of" means that any one or more of these features may be provided individually or in combination. For example, "at least one of [A], [B], and [C]" covers any of the following options: A alone (without B or C), B alone (without A or C), C alone (without A or B), a combination of A and B (without C), a combination of A and C (without B), a combination of B and C (without A), or a combination of A, B, and C.

[0143] While the present invention has been described in detail herein with reference to illustrative embodiments thereof, it is to be understood that the invention is not limited to those precise embodiments, and that various changes and modifications may be effected therein by those skilled in the art without departing from the scope of the invention as defined by the appended claims.

Claims

1. An apparatus, the apparatus comprising: a memory access circuit configured to perform stack access to a stack data structure based on a stack pointer; and a processing circuit configured to perform a stack pointer switch validity check operation associated with a switch of the stack pointer from an outgoing stack pointer value to an incoming stack pointer value, the stack pointer switch validity check operation including verifying whether an incoming data value obtained by the memory access circuit in response to a memory access request specifying an address determined based on the incoming stack pointer value conforms to at least one stack capping value validity condition, the at least one stack capping value validity condition including a condition that a predetermined portion of the incoming data value corresponds to a given paging address indicating paging of an address space including the address determined based on the incoming stack pointer value, wherein: the processing circuit is configured to determine whether the incoming data value conforms to the at least one stack capping value validity condition, regardless of whether another portion of the incoming data value other than the predetermined portion corresponds to a sub-paging address bit of the address determined based on the incoming stack pointer value; and the processing circuit is configured to trigger an error handling response in response to determining that the incoming data value does not conform to the at least one stack capping value validity condition.

2. The apparatus according to claim 1, wherein the given paging address is a virtual paging address indicating paging of a virtual address space including a virtual address determined based on the incoming stack pointer value.

3. The apparatus according to any one of claims 1 and 2, wherein the processing circuit is configured to perform an outgoing stack capping operation associated with the switch of the stack pointer from the outgoing stack pointer value to the incoming stack pointer value to push a valid stack capping value into a location having an address selected based on the outgoing stack pointer value, the valid stack capping value specifying in the predetermined portion thereof a paging address indicating paging of an address space including the address selected based on the outgoing stack pointer value.

4. The apparatus according to claim 3, wherein the paging address specified in the valid stack capping value is a virtual paging address indicating paging of a virtual address space including a virtual address selected based on the outgoing stack pointer value.

5. The apparatus according to any of the preceding claims, wherein the stack pointer is a guard control stack (GCS) pointer for controlling access to a guard control stack (GCS) data structure for protecting a return address upon return from an exception or a function call.

6. The apparatus according to claim 5, wherein the apparatus includes a memory management circuit configured to determine whether to permit access to a target address based on memory attribute data associated with the target address, the memory attribute data specifying whether a target memory address space region containing the target address is a GCS region for storing the GCS data structure, wherein write access to the GCS region is restricted to a dedicated class of GCS access instructions.

7. The apparatus according to claim 6, wherein when the memory attribute data specifies that the target memory address space region is the GCS region, the memory management circuit is configured to reject instruction fetch or a branch to the target address.

8. The apparatus according to any one of claims 6 and 7, wherein the memory management circuit is configured to, in response to determining that the target memory address space region is the GCS region, reject a non-GCS store operation to the target address triggered by a store instruction other than the dedicated class of GCS access instructions.

9. The apparatus according to any one of claims 6 to 8, wherein the memory management circuit is configured to, in response to determining that the target memory address space region is not the GCS region, reject a GCS load / store operation to the target address triggered by one of the dedicated class of GCS access instructions.

10. The apparatus according to any of the preceding claims, wherein the at least one stack capping value validity condition further includes a condition that the least significant portion of the incoming data value has a bit pattern that cannot be specified by the least significant portion of any valid instruction address.

11. The apparatus according to the preceding claim, wherein the processing circuit is configured to perform the stack pointer switch validity check operation in response to a first stack pointer switch instruction specifying an operand indicating the incoming stack pointer value.

12. The apparatus according to claim 11, wherein the apparatus includes a stack pointer register configured to store the stack pointer; wherein: in response to the first stack pointer switch instruction, when the stack pointer switch validity check operation is successful, the processing circuit is configured to update the stack pointer register from the outgoing stack pointer value to the incoming stack pointer value.

13. The apparatus according to claim 12, wherein in response to the first stack pointer switch instruction, the processing circuit is configured to push an in-progress token value specifying the outgoing stack pointer value to a location having an address selected based on the incoming stack pointer value.

14. The apparatus according to claim 13, wherein for a load operation for loading the incoming data value to be verified in the stack pointer switch validity check operation and a store operation for pushing the in-progress token value to the location having the address selected based on the incoming stack pointer value, the processing circuit is configured to perform the load operation and the store operation atomically.

15. The apparatus according to any one of claims 13 and 14, wherein in response to a second stack pointer switching instruction, the processing circuit is configured to: Verify whether a given data value obtained by the memory access circuit in response to a memory access request specifying an address determined based on the current stack pointer in the stack pointer register is a validly formed in - progress token value; and In response to verifying that the given data value is a validly formed in - progress token value: Trigger a write of a valid stack cap value to a location having an address determined based on a given stack pointer value specified by a part of the given data value, the valid stack cap value specifying in a predetermined part of the valid stack cap value a paging address indicating paging of an address space including the address determined based on the given stack pointer value.

16. The apparatus according to claim 15, wherein in response to the second stack pointer switching instruction, when the given data value is verified to be the validly formed in - progress token value, the processing circuit is configured to update the current stack pointer in the stack pointer register to indicate removal of the entry that will provide the given data value from the corresponding stack data structure.

17. The apparatus according to any one of claims 15 and 16, wherein when the least significant part of the given data value has a bit pattern that cannot be specified by the least significant part of any valid instruction address and cannot be specified by any value that meets the at least one stack cap value validity condition, the processing circuit is configured to verify that the given data value is the validly formed in - progress token value.

18. The apparatus according to any one of claims 15 to 17, wherein the processing circuit is configured to trigger an error handling response in response to determining that the given data value is not a validly formed in - progress token value.

19. The apparatus according to any of the preceding claims, wherein the error handling response includes at least one of the following: Signaling an error; Setting an error reporting indication; and Setting the stack pointer to an invalid value.

20. A method, the method comprises: Performing a stack pointer switching validity check operation associated with a switch of a stack pointer from an outgoing stack pointer value to an incoming stack pointer value, the stack pointer switching validity check operation including verifying whether an incoming data value obtained by a memory access circuit in response to a memory access request specifying an address determined based on the incoming stack pointer value meets at least one stack cap value validity condition, the at least one stack cap value validity condition including a condition that a predetermined part of the incoming data value corresponds to a given paging address indicating paging of an address space including the address determined based on the incoming stack pointer value, wherein: Whether the incoming data value meets the at least one stack cap value validity condition is determined regardless of whether another part of the incoming data value other than the predetermined part corresponds to a sub - paging address bit of the address determined based on the incoming stack pointer value; and The method includes triggering an error handling response in response to determining that the incoming data value does not meet the at least one stack capping value validity condition.

21. A computer program comprising instructions that, when executed by a host data processing device, control the host data processing device to provide an instruction execution environment for executing target code. The computer program comprises: memory access program logic for performing stack access to a stack data structure based on a stack pointer; and handler logic for performing a stack pointer switch validity check operation associated with a switch of the stack pointer from an outgoing stack pointer value to an incoming stack pointer value. The stack pointer switch validity check operation includes verifying whether an incoming data value obtained by the memory access program logic in response to a request specifying an address determined based on the incoming stack pointer value meets at least one stack capping value validity condition. The at least one stack capping value validity condition includes a condition that a predetermined portion of the incoming data value corresponds to a given paging address indicating paging of an address space including the address determined based on the incoming stack pointer value, wherein: the handler logic is configured to determine whether the incoming data value meets the at least one stack capping value validity condition, regardless of whether another portion of the incoming data value other than the predetermined portion corresponds to a sub-paging address bit of the address determined based on the incoming stack pointer value; and the handler logic is configured to trigger an error handling response in response to determining that the incoming data value does not meet the at least one stack capping value validity condition.

22. A storage medium storing the computer program according to claim 21.