Processor with hardware-supported memory buffer overflow detection - Patents.com
The processor's hardware-based buffer overflow protection circuitry with write protection indicators addresses inefficiencies and vulnerabilities in conventional stack overflow prevention, enhancing security and reducing development burdens.
Patent Information
- Application Number
- JP2020573096
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-03-16
- Filing Date
- 2019-03-18
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2039-03-18
AI Technical Summary
Conventional software and hardware techniques to prevent stack overflows are inefficient, costly, require extensive development efforts, and can be bypassed by malicious attacks like Return Oriented Programming (ROP), leading to security breaches and unintended program execution.
A processor with hardware-based buffer overflow protection (BOVP) circuitry that includes a memory stack with write protection indicators, detecting and preventing unauthorized writes to stack storage spaces by generating faults when write protection is violated.
The hardware solution effectively safeguards against accidental and malicious stack overflows, reducing development overhead and enhancing security by ensuring the integrity of program execution flow.
Smart Images

Figure 0007679569000003 
Figure 0007679569000004 
Figure 0007679569000005
Abstract
Description
[Technical field]
[0001] This application relates generally to processors (such as microprocessors, digital signal processors, and microcontrollers), and more particularly to processors with hardware support for protecting against unwanted memory buffer overflows. [Background technology]
[0002] A processor includes or has access to a portion of memory that is typically referred to as a stack, which may also be referred to as a call stack, execution stack, program stack, and other additional notations. The term "stack" is based in part on the fact that the memory portion is last-in, first-out, so that as information is added to memory, the information is aggregated with the existing information already there to form the stack of additional information, and then as the information is removed from memory, the stack of aggregated information is decreased. Also, in this context, information added to the stack is typically shown to be at the "top" of the stack for purposes of illustration, although in practice in some architectures the memory area that includes the stack may be addressed using increasing memory addresses, and in other architectures decreasing memory addresses are used. Even in the latter case, the lowest address is considered to be the "top" of the stack.
[0003] The primary type of information stored in the stack is addresses that represent points in a sequence of executable programming code. More specifically, typically, as code is being executed by a processor, a memory store such as a register, often called a program counter, stores the memory address of each of the instructions currently being executed. The program counter is so named because code is generally executed sequentially. The program counter can count, or increment, thereby proceeding to execute the next instruction at the address immediately following the address of the previously executed instruction, and similarly for a particular block of code. However, at various points in this type of sequential addressing of executable instructions, a change in the sequence may be desired, which is accomplished by sequence-changing instructions, which may include "call," "branch," "jump," or other instructions, thereby directing the execution sequence to a target instruction whose address is not the next sequential address following the currently executed instruction. The target instruction may be part of a group of other instructions, sometimes called a routine or subroutine. In connection with sequence-changing (e.g., call) instructions that result in changing addressing to something other than continuing in a sequential manner, one common approach is for the current or next increment value of the program counter to be stored (often referred to as "pushed") onto a stack, so that after the target routine is completed, address flow is returned to the sequence that occurred before the call. That is, the instruction sequence "returns" to the next instruction following the call to that routine. Alternatively, some architectures (such as Advanced RISC Machines (ARM)) do not immediately push the program counter onto the stack; instead, the return address is stored in a register only if the called function calls (or may call) another function, and the value of the return register is then pushed onto the stack.In any case, once an address has been pushed onto the stack, a return following the call may be accomplished by taking the address that was pushed onto the stack when the call occurred, and that value is said to be popped from the stack, thereby removing it and causing the top of the stack to be moved to the next lowest word on the stack. The above description assumes only a single call to a routine, which ultimately causes the routine to return directly to the address that followed the sequence following the single call. However, the first routine may call a second routine while it is active and prior to its completion, in which case this latter call will again push an additional return address onto the stack. The additional program counter address identifies the executable instruction address to which program flow should return upon completion of the second routine that was called. Thus, in the example of two consecutive calls (prior to the return), there may be on the stack the program return address when the first routine was called, and above that the program return address when the second routine was called. This process may be repeated for multiple routines, so that the stack receives an additional new address for each additional call, and so that each additional address is successively popped as a return is executed from each successively called routine. Thus, the stack provides an indication to the eventual completion of each called routine and the return execution address as each respective routine is called.
[0004] As explained above, stack technology has long provided a sound method of controlling executable instruction flow, but an unfortunate by-product has been unintentional defects of, or deliberate attacks against, the stack when program counter addresses in the stack are overwritten before they are needed to re-establish proper executable code flow. For example, a stack typically has a maximum capacity provided by a finite number of storage locations. Thus, when an address is written beyond the maximum capacity, a stack overflow is said to have occurred, and erratic behavior or reported defects occur when the event is made known, for example, via software running on the system (e.g., an operating system) or the like, or via a "top of stack register." The "top of stack register" is used in some hardware approaches to indicate the top stack location, and therefore can also be used to detect when that location has been exceeded. As another example, in addition to the program counter address, in some architectures the stack is used to store data, usually temporarily. Such storage locations are typically called "stack frames" or "call frames." Such a frame is usually architecture and / or Application Binary Interface (ABI) specific, but it often contains the parameters and return address for the called function, along with any temporary data space for the called function. Thus, in this case, an additional memory location in the stack is temporarily reserved for that stack frame, which is near or includes the stored return address, and is within the stack memory space. If the buffer is filled beyond its intended size, such filling may overwrite valid return addresses and / or data stored in the current or previous stack frames.
[0005] In addition to the above, various nasty attack techniques have evolved to gain control of a processor and impair its intended operation, with the aim of "hacking" or otherwise interfering with a computing system, resulting in a variety of outcomes that may include anything from relatively good operation to damage to important functionality and security breaches. In this regard, stack overflow is a common tool of such attacks, sometimes referred to as "smashing the stack". In this case, the "hacker" attempts to move program control to a location that is not a valid operational stack return address. This may be accomplished, for example, by overwriting a valid return address with a different, illegal address, whereby the program flow, upon returning to the stack, pops the illegal address and directs the executable flow to other instructions. Such other instructions may also be maliciously loaded into the system and thereby executed following the smash. Alternatively, hackers have also learned to use subsets or excerpts of valid existing code, sometimes referred to as gadgets, where by stitching together different gadgets in sequence, further nasty results may be reached. The respective functionalities of each gadget are then combined to achieve an unintended function or result that the hacker implemented in the system that was originally designed. This latter approach is sometimes called Return Oriented Programming (ROP) because each gadget ends with its own return. The return at the end of each gadget pops the return address (i.e., the alternate address written by the hacker via the buffer overflow) from the stack. This causes the processor to continue execution, starting after that alternate address, which in turn initiates the start of another gadget, which can then be repeated with multiple such gadgets. In fact, the use of gadgets in this problem allows the hacker to approach or reach the so-called Turing (named after the mathematician Alan Turing) completeness.This generally means providing sufficient data manipulation operations to simulate any single-taped Turing machine, i.e., to accomplish all the algorithms in the set of such algorithms, which in a loose interpretation means the ability to accomplish a wide range of functions.
[0006] As described above, conventional software and hardware (e.g., Data Execution Prevention (DEP)) techniques attempt to prevent and / or detect unintended and / or malicious instruction changes of program control flow via overwritten return addresses in a buffer / stack. However, such approaches have one or more drawbacks, including (a) high overhead costs, (b) requiring developers to implement or learn new tools, (c) extensive testing, (d) recompilation and / or source code modifications are required, which are not always possible for third-party libraries, and (e) hackers have found ways around conventional approaches, including ROP, as workarounds for DEP. Thus, while conventional software mitigation techniques for stack overflow serve various needs, improvements are possible. Summary of the Invention
[0007] In an exemplary embodiment, a processor includes: (a) circuitry for executing program instructions, each program instruction stored at a location having a respective program address; (b) circuitry for storing an indication of the program address of the program instruction to be executed; (c) a memory stack including a plurality of stack storage spaces, each stack storage space operable to receive a program address as a return address; (d) a plurality of write protection indicators corresponding to each stack storage space; and (e) circuitry for generating a fault in response to detecting a processor write directed to a selected stack storage space, wherein for the selected stack storage space, a respective write protection indicator in the plurality of write protection indicators indicates that the selected stack storage space is write protected. [Brief description of the drawings]
[0008] [Figure 1] FIG. 2 illustrates an electrical and functional block diagram of a processor of an exemplary embodiment.
[0009] [Diagram 2] 2 shows a more detailed block diagram of the BOVP circuit 18BOVP and stack 24STK of FIG.
[0010] [Diagram 3] 3 shows a flow chart of an example embodiment of a method 300 of operation of the processor 10 of FIG.
[0011] [Figure 4a] Stack 24STK is shown in an application of method 300 of FIG. [Figure 4b] Stack 24STK is shown in application of method 300 of FIG. [Figure 4c] Stack 24STK is shown in an application of method 300 of FIG. [Figure 4d] Stack 24STK is shown in an application of method 300 of FIG. [Figure 4e] Stack 24STK is shown in an application of method 300 of FIG. [Figure 4f] Stack 24STK is shown in an application of method 300 of FIG.
[0012] [Diagram 5] The block diagram of FIG. 2 is shown, but with an alternative embodiment for the stack.
[0013] [Figure 6] 3 shows an electrical and functional block diagram of another embodiment of the implementation of the internal memory 24 and WP management block 32 of FIG. 2. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] 1 illustrates an electrical and functional block diagram of an exemplary embodiment processor 10. The processor is one of various computing circuits that include the functions described herein, and thus may include any of a microprocessor, digital signal processor, microcontroller, etc. Processor 10 includes various conventional blocks and improvements of the exemplary embodiment, each of which is described hereinafter.
[0015] Processor 10 is typically implemented as a single integrated circuit device 12, as indicated by the appropriate blocks in Figure 1. Within that block (or individually, as may be partitioned), processor 10 generally includes three main blocks, including a control unit 14, an algorithmic logic unit (ALU) 16, and a memory unit 18, each of which may communicate with each other by bidirectional coupling to bus B, as shown in Figure 1. Bus B is also connected to receive data from an external input device 20 and provide data to an external output device 22. Memory unit 18 is also connected to bus B. M1 Internal memory 24 and bus B M0The control unit 14 may communicate with an external memory 26 via a program counter 14 which may be implemented elsewhere in the processor 10. More particularly, the control unit 14 controls the initiation and operation of the other blocks and generally initiates sequences of instructions of executable code, usually in response to any higher level programming, including possibly source code and / or operating system software. In this regard, the control unit 14 therefore also controls a program counter 14 which may be implemented elsewhere in the processor 10. PC 1. The control unit 14 is illustrated to include a generic representation of a processor 10. Thus, the control unit 14 may cause the memory unit 18 to fetch program instructions from either an internal memory 24 (e.g., RAM, ROM) or an external memory 26 (e.g., Flash) according to a sequence defined by a particular program, and potentially using other efficiency measures (e.g., prediction techniques, etc.). The instructions fetched by the memory unit 18 are provided to one or more execution pipelines associated with the ALU 16, where they are executed, again under the control of the control unit 14, to accomplish data movement, as well as logical and mathematical operations, the results of which are again returned to the memory unit 18 for storage in either the internal memory 24 or the external memory 26. In a broad sense, the input / output of instructions and data described above may also be accomplished in combination with receiving such information from the input device 20 and providing such information to the output device 22. Finally, this description of the processor 10 is a broad overview of a processor architecture. However, a processor may include or be characterized as including various other aspects related to processing architecture, connectivity, and functionality.
[0016] 1, the following description provides an introduction to an improved example embodiment of the processor 10. Specifically, the internal memory 24 includes a memory stack 24, which is a dedicated portion of the memory space in the internal memory 24. STKwhere the number of words stored in the stack can vary widely depending on the architecture, application binary interface (ABI), and the particular application code. Also, the word width can vary depending on the CPU architecture (e.g., 2×16, 32, 64, etc., for N≧4). N Also, the stack is defined by 24 STK Generally speaking, the program counter l4 PC From or Program Counter l4 PC and to receive addresses pushed (directly or via an intermediate register) relative to the program counter l4 from the stack, either directly or via a return register (e.g., the link register in an ARM processor). PC The stack 24 exists and functions to facilitate popping back to the stack, thereby returning the program execution flow to the next instruction following the instruction that generated the pushed address, while providing some buffer space in which other information may be temporarily stored and easily accessed from the stack memory space. STK Only one stack is shown because a system may contain one or more such stacks, where each stack is a region of memory and, in many systems, may begin or end anywhere. For example, multiple stacks are common when using a real-time operating system (RTOS), where each thread typically has its own stack.
[0017] At the end of FIG. 1, for the exemplary embodiment, stack 24 STK (as well as other stacks) are buffer overflow protection (BOVP) circuitry 18 shown as part of memory unit 18. BOVP In particular, as described below, the BOVP circuit 18 BOVP and Stack 24 STK Especially in combination with Stack 24 STKThis provides hardened hardware protection against direct or accidental overwrites of the return address on the stack, which can occur by mistake (e.g., incorrect use of certain code, such as in C or C++), or by deliberate hacking attacks aimed at deliberately smashing the stack.
[0018] Figure 2 shows the BOVP circuit l8 of Figure 1. BOVP and Stack 24 STK 1 shows a more detailed block diagram of the BOVP circuit 18. BOVP The BOVP circuit 18 includes a read / write (R / W) detection block 30 and a write protection (WP) management block 32. Each of these additional blocks is described below and may be configured as described herein. BOVP is preferably a hardware implementation, since embodying even a portion of it in software would impose significant, if not prohibitive, runtime overhead.
[0019] BOVP Circuit l8 BOVP 2. The stack 24 is connected to appropriate circuitry, which may be based on the particular platform on which the processor 10 (e.g., ARM, x86, and one of the various processors manufactured by Texas Instruments) is implemented. STK Each instance in which a read or write is requested in relation to the stack is detected. For example, a write may occur in relation to a push instruction that attempts to push a return address onto the stack. Similarly, a write may occur when data is copied from one location to the stack, and so a write may occur in relation to the stack 24. STK As another example, a read may occur in conjunction with a pop instruction, e.g., a portion of stack 24 may be temporarily allocated as a buffer. STK As a final example, the program tries to read a return address that was previously pushed onto the stack, which is temporarily allocated as a buffer. STK A read can also occur when data is read from a portion of the.
[0020] WP management block 32 stack 24 STK In the exemplary embodiment shown in FIG. 2, the write protection data is stored in the stack 24. STK Thus, in FIG. 2, the stack 24 STK are each illustrated as having a 32-bit memory word, and appended to each word (e.g., next to the least significant bit) is an additional write protect WP indicator bit. Thus, in one exemplary embodiment, stack 24 STK In one embodiment, the memory (e.g., SRAM) for containing the WP indicator bit is configured to include one additional bit beyond the nominal size of the memory word (e.g., 32 bits). In such an approach, the processor memory may be subdivided so that one portion of the memory is only word-wide and another portion includes the additional bit per memory word. As described below, alternative exemplary embodiments also contemplate associating a WP indicator bit with each memory word without expanding the width of the memory word beyond the nominal word size. In any event, as described further below, the WP management block 32 manages the WP indicator bit in the stack 24. STK 2, and in this description, each word (and its associated WP indicator bit) is separately addressable, as indicated by the convention in FIG. <xx>where xx is an address or location indicator. For example, in FIG. 2, 32 such addresses are <00> 2 illustrates a stack pointer SPTR, which is a value (e.g., a value stored in a register) that identifies the current stack address and, possibly, its corresponding WP indicator bit. Thus, in FIG. 2, the stack pointer SPTR indicates the location of an empty stack 24 STK The top of the STKA <00> ), i.e., points to the next location where a return address can be pushed (here, the top is the first word that can be written, since no information has been written yet). Also in the example of FIG. 2, the stack pointer SPTR is shown moving toward each new top of the stack by moving vertically upward with each respective increment of the stack address. However, it is contemplated that some stack orientations have a "toward the top" stack pointer movement associated with a decrement of the stack address. It is contemplated that in any event in this description, adding a return address to the stack effectively changes the top of the stack. As indicated by the stack pointer, the addition of a return address advances the stack to the next location where information can be added (hereafter referred to as the "top" of the stack). Such movement may be accomplished via incrementing or decrementing the stack address in different exemplary embodiments.
[0021] 3 illustrates a flow chart of an exemplary embodiment of a method 300 of operation of processor 10. The flow chart makes particular reference to, and is implemented by, the blocks of FIG. 2, all of which are implemented by stack 24. STK Thus, the method 300 relates to reading and writing the stack 24. STK The method begins with a stack access step 302, where R / W detection 30 detects when either a read or a write is requested on a stack. In this regard, R / W detection 30 includes sufficient circuitry (e.g., logic) coupled to appropriate signals, such as an instruction bus, trace interface signals, or signals from other hardware, that can be monitored and detected when a stack read / write operation is attempted, while also distinguishing stack pushes from other stack writes, and stack pops from other stack reads. Identifying, monitoring, and detecting such signals can be accomplished given the particular architecture of processor 10. In any event, if the requested operation is a write, method 300 continues to step 304, and if the requested operation is a read, method 300 continues to step 306. Each of these options is described below.
[0022] Write protection according to the use of an exemplary embodiment of the WP indicator will now be introduced and described below. STK In order to protect each return address pushed onto the stack from being subsequently overwritten, a WP indicator bit is set. Thus, if an attempt is made to overwrite a return address stored on the stack that is protected by the respective WP indicator bit, a fault is generated to prevent the overwriting of the existing, protected return address that is already stored. Thus, if a stack access is made to write either a new return address or buffer information to a stack location that does not previously have an existing, protected return address stored therein, the write access is permitted. Also, if the stack access made is to write a new return address to a stack word that is not protected (i.e., the WP indicator is not set), the write is used to set the WP indicator to subsequently protect the new return address from being overwritten. However, if a stack access is made to push a new return address onto a stack location that already stores an existing, protected return address, a fault is generated to prevent the overwriting of the existing, protected return address that is already stored. These and other contingencies are further described below.
[0023] Stack 24 STK In step 304, which is reached because a detected write access to an address has been detected, step 304 evaluates the WP indicator bit for the address to which a write is being attempted to store data, and in particular evaluates whether the WP indicator bit is already set. If the WP indicator bit is set, method 300 proceeds from step 304 to step 308, whereas if the WP indicator bit is not set, method 300 proceeds from step 304 to step 310.
[0024] Step 308, reached because a write was attempted to a stack address with a corresponding set WP indicator bit, generates a fault, thereby reporting the fault. Thus, in FIG. 2, this event is indicated by an asserted FAULT signal coupled to bus B, and may be received and responded to by any circuitry coupled to bus B. Fault reporting may be accomplished in a variety of ways, such as via an interrupt to any further program execution, actual or speculative, and then some event handler processes the fault accordingly. The fault setting of step 308 also asserts a fault when a push or non-push write is attempted to a stack location whose corresponding WP data is set (e.g., as the 33rd bit in an SRAM word). Also, as introduced above, if the return address is set to stack 24, then the return address is set to stack 24. STK , the WP indicator bit is set, so steps 304 and 308 generate a fault if there is an attempt to overwrite such a pushed return address. In this manner, whether such an attempt is by mistake (e.g., erroneous C code) or a deliberate attempt to corrupt the stack (e.g., a ROP attack), a flaw or undesired change in outcome, control, and / or security that may result from allowing the return address to be overwritten is avoided, and instead a fault is generated. Thus, although not shown in FIG. 3, method 300 may, for example, branch to an appropriate available fault handler based on the response to the fault, and then return to method 300.
[0025] Step 310 is reached because the requested stack access is a write operation and the WP indicator is not set for the address to be written to. <stka>, performs the requested write as a push read, or as a buffer write, e.g., in the case of a non-push write. As a push, the write is performed from another register or from the program counter l4. PC , so that upon return, the next processor executed instruction is at the return address. Next, method 300 continues to step 312.
[0026] Step 312 is a step in which the write operation of step 310 is performed on stack 24. STK , which includes attempting to write a return address to a stack (e.g., stack 24). Often, a push operation, by definition, always pushes a stack STK ), but which are considered herein to be possible push operations that do not push a return address. Thus, also included in step 312 is a determination of whether the push is a return address. This determination may be accomplished in example embodiments by, for example, detecting when the value comes from the PC or a register that is specifically used to hold a return address (e.g., the LR register in the ARM architecture). Accordingly, some example embodiments monitor processor signals to determine whether this is the case. On the other hand, in an ARM implementation of example embodiments, example embodiments rely on a compiler convention that the first instruction executed in a function is a push of a value that includes the LR register value. If step 312 detects that the requested write is a push operation, method 300 continues with step 314, whereas if step 310 detects that the requested write is not a push operation, method 300 returns to the next instance of step 302.
[0027] In step 314, which is reached because the detected write to stack operation was also detected to be a push operation (which attempts to write an address to the stack), the WP management block 32 then sets the WP indicator bit for the stack memory location into which the return address is to be written, i.e., the location of the current stack address pointer. <stka>and the WP indicator for that position is WP[SPTR <stka>2.]. Thus, as used herein, setting this indicator bit may be by asserting it to a binary value of 1, thereby indicating the return address STKA being written to the same respective address. As long as the WP indicator remains set, the STKA is write protected. Also, when this instance of step 314 occurs with respect to the point shown in FIG. 2, the set WP indicator indicates that the SPTR <00> to a binary 1. Also, in other example embodiments, optionally, instructions may be permitted to set any WP bit, thereby allowing protection for any corresponding address on the stack. In any event, after step 314 sets the WP indicator, method 300 returns to the next instance of step 302.
[0028] From the above discussion, step 310 is reached after a write access has been triggered, either as a push or some other type of write (e.g., a buffer copy). And so, upon completion of steps 312 and 314, for a given stack address, a valid return address has been written and its respective WP indicator bit has been set. Of course, in the case of a non-push write, the data is also written to location SPTR. <stka>, but the WP bit is not set for such written data. Also, in different example embodiments, the order of steps 314 and 312 is reversed, or they could occur simultaneously as an atomic step, provided that both occur reliably if the write to the stack is a return address. From an implementation standpoint, one may be easier than the other, depending on the implementation details. In any event, once both steps (or just step 312, in the case of a stack write that does not include a return address) are completed, method 300 returns from step 312 to step 302, thereby awaiting the next stack access.
[0029] Having described the write-related steps of method 300 (shown on the right side of FIG. 3), attention is now directed to the method steps when step 302 determines that the stack access is a read, beginning with step 306. Step 306 determines whether the read is a pop operation. A pop operation is one in which a stack 24 is popped. STK The first step of the method 300 is to read the return address from and place it in the CPU program counter l4PC, thereby restoring the program flow to the next address following the address of the instruction (e.g., a branch) that left the previous program sequence. The determination of step 306 may also be made in response to detecting a signal, based on the processor architecture. If step 306 detects that the requested read is a pop operation, the method 300 continues with step 322. On the other hand, if step 306 detects that the requested read is not a pop operation, step 322 is skipped and the method 300 continues with step 324. The detection of a pop operation may be made by monitoring the value at the stack pointer. When the stack pointer value is moved towards the bottom of the stack, all values between the previous stack pointer value and the new stack pointer value are considered to have been popped. In step 322, which is reached because the detected read from the stack operation is further detected as a pop operation, the WP management block 32 checks the stack memory location from which the return address is to be read, i.e., the current stack address pointer (i.e., WP[SPTR <stka>]) location. Thus, in this specification, clearing of this bit is accomplished by deasserting the bit to a binary value of 0. With the WP indicator thus deasserted to a value of 0, if the address STKA is later deemed to be written to, step 304 described above related to stack writes may detect the cleared WP indicator, thereby allowing the write by transitioning to step 312. In any event, after step 322, method 300 continues from step 322 to step 324.
[0030] In step 324, the requested read is performed, either as a pop read after step 322, or as a direct buffer read from finding a non-pop read in step 306. Thus, as a pop, the read is a read of the return address (or other information) from the stack, along with the current stack address pointer (i.e., SPTR <stka>) of the memory word at location SPTR. Although not shown, the read could be to another register or directly to the program counter 14Pc, so that the next instruction the processor executes will be the instruction at the return address. Thus, upon completion of steps 322 and 324, for a given stack address, a valid return address has been read and its respective WP indicator bit has been cleared. Of course, in the case of a non-pop read, the data is also stored at location SPTR. <stka>The stack is read from and provided to the destination as may be indicated by a non-pop read instruction. Also, similar to the above description of steps 314 and 312, the order of steps 322 and 324 may be reversed, provided that if a read from the stack involves a pop, then both occur reliably, and implementation may depend on other implementation details. In any event, once steps 322 and 324 are completed, method 300 then returns from step 324 to step 302, thereby awaiting the next stack access.
[0031] 4a to 4f show stack 24 STK 3 illustrates an example sequence of the method 300 using the representation in various memory words and the movement of the stack pointer SPTR. In this first example, it is assumed that a programming language (e.g., C) implements according to the code shown in Table 1 below. [Table 1] Before any instruction is executed, the stack STK is expressed similarly to the case in Figure 2, and the stack pointer SPTR is the <00> )
[0032] Next, the first instruction calls the "hello" routine according to Table 1. In that case, the corresponding assembly language would write the program counter address stored in the LR register onto the stack 24. STK The attempted push is thus a stack access write attempt, as detected in step 302 of method 300. Also, the attempted write is to an empty stack location with a cleared WP indicator, as determined in step 304, so flow proceeds to step 310. Step 310 further identifies the write as a return address push (e.g., a push of LR), which is passed to step 314 to check for the STK. <00> This is followed by step 312, where the return address is pushed onto that same address. In this regard, FIG. 4b shows the stack 24 after these steps. STK In this case, step 314 is <00> The WP indicator bit for the STKA is set to a value of 1, and step 312 transfers the LR data to the STKA <00> is writing.
[0033] Next, the result of the C instruction "uint32t buf[4]" is stack 24. STK In this regard, the instruction "uint32t buf[4]" causes the stack pointer SPTR to be appended with a value of 4. It therefore reserves buffer space on top of the pushed LR value in Figure 4b. <00> From STKA in Fig. 4c <04> As shown in FIG. 4c,
[0034] Next, the C instruction "memcpy(buf,rx_buf,len);" reads a buffer of words with length set by the variable rx_buf; from another memory to the stack. STK So in the current example, the rx_buf is defined to have 4 words, so each of those 4 words is copied to the stack 24. STK Thus, each write represents an attempt to access the stack, which is again detected by step 302 of method 300. For example, for the first write, step 310 first accesses the STKA address in question (i.e., STKA <04> ) is set. In the current example, the WP data is not set, so the next step 304 determines that the write is not a push, and then step 312 performs the write, as shown in FIG. 4d. STK The same process is repeated for each of the remaining three words to be copied to the stack 24, so that upon completion of three more iterations of steps 304, 310, and 312, the stack 24 STK is shown in Figure 4e. Also, in this context, the address used for rx_buff accesses is tracked separately from the stack pointer, so the stack pointer does not move with each write to the buffer, and therefore SPTR is the address of STKA. <04> It is shown as pointing to the.
[0035] As mentioned above, the buffer data is stored in the stack 24 STK The stack pointer SPTR has been successfully stored at address STKA and may be quickly and easily accessed as desired during the routine. Eventually, the routine completes, as indicated by the C language "return;" instruction, as shown in Table 1. In response to that instruction, a value of -4 is added to the stack pointer SPTR, thereby assigning it to address STKA. <00> , thus pointing to the LR return address, as shown in FIG. 4f. Then, an assembly language POP instruction is used to call a read from that address, again invoking steps 306, 322, and 324 of method 300. Thus, in step 322, the previously set WP indicator bit is cleared, and in step 324, address STKA is placed in the LR return address, as shown in FIG. 4f. <00> The LR value in is read, and from there, that value is set to the processor's program counter l4. PC so that the next executed instruction is at the return address, thereby properly returning execution to the instruction following the routine call.
[0036] The above example of FIG. 4a to FIG. 4f shows the return address stack 24 STK 4d and 4e, the above examples assume that the buffer length is four words, as passed by the variable rx_buff; However, we now assume that the variable rx_buff is five or more words (e.g., eight words, in which case rx_buff[8]) either by mistake or by a deliberate hacking operation. Referring to FIG. 4e, each of the first four words is pushed to stack 24, and the first four words are pushed to stack 24. STK After being copied to STKA, the next write attempt, i.e., rx_buff[5], <00> This attempted write is first detected by step 302 of method 300, and then step 304 detects the STKA <00> 4e, the condition at step 304 is affirmative, i.e., the WP indicator bit is set for that location. As a result, method 300 continues to step 308, where a fault is generated. Also, as part of or in addition to the fault generation, writes are blocked in response to the WP indicator bit being set. Thus, in an exemplary embodiment, part of the fault generation (or in addition to or instead of the fault generation) is to generate a fault for the STKA. <00> This thwarts the C language's efforts to overwrite the stored return address in , thereby avoiding "smashing" the stack. Thus, STK If there is a subsequent return to , the proper return address is maintained, and there is no ability to jump or transition to an address other than the return address of the regular stack location. Admittedly, the above example of copying a buffer is just one possible approach, but in all events, earlier detection of the push of the return address onto the stack is protected from subsequent overwrite by each set WP indicator bit, provided there is no subsequent pop of that address, and until there is a subsequent pop, in which case each set WP indicator bit is cleared.
[0037] FIG. 5 shows the block of FIG. 2 again, but with the reference added to the stack 24' STK Stack 24 is an alternative exemplary embodiment of the stack. STK Similarly, 24' STK includes a WP indicator bit for each address location where address data may be written, but rather than appending an additional bit to each word, the set of WP indicator bits is stored in a single memory location, which may be a memory location somewhere near the stack, so for example, the stack address <00> , and is therefore denoted as address <00-2>. Alternatively, in another example, a single WP store word may be one of the stack words, where the additional word itself may or may not be protected. Also, stack 24' STK Each memory word in STK As in the case of STK In addition, since a single memory word can thus support a total of 32 WP indicator bits, each different WP indicator bit preferably corresponds to a respective different memory location in the stack 24'STK. <00> The WP indicator bit for the lowest stack word in the <00> Each higher bit in <00> The WP indicator bits in STKA<00-2> represent the WP indicator bit for the next top stack word above. Each WP indicator bit in STKA<00-2> maps to a respective word as shown in Table 2 below. [Table 2] Finally, in FIG. 5, the BOVP circuit 18 BOVP WP management block 32 is shown coupled to access only the word at address <00-2> because that is the word that reads / writes with respect to the WP indicator bit. Thus, consistent with various steps in method 300, WP management block 32 may read or write each bit of the word at address <00-2>, such as by reading the entire word and then reading or writing bits therein (and in the latter case subsequently writing the entire word, including the modified bits, back to address <00-2>).
[0038] 6 shows an electrical and functional block diagram of another exemplary embodiment implementation of the internal memory 24 and the WP management block 32. Each of these blocks is described below.
[0039] Looking first at internal memory 24, a stack 24" may be used to denote that a device (e.g., a processor) may include one or more stacks in its memory space. STKS is shown. Thus, while FIG. 5 illustrates a single block of R / W memory space containing 32 address / buffer words and one WP indicator word, FIG. 6 provides multiple (e.g., three) stack blocks with each block having its own WP indicator space. Although no particular size is provided in the illustrated approach, in alternative example embodiments, the stack size may vary, and thus the WP bit corresponding to each stack may vary. Thus, in one example, each stack may be 32 words and each WP bit space may be 32 bits, but in other variations, the stack sizes may all be the same and other than 32 words. Alternatively, one or more may be different from the others, and thus the space required to accommodate each WP bit may differ. Thus, in this regard, the generality of FIG. 6 is intended to illustrate the breadth of alternative example embodiments, and the WP bit (or other indicator) for a given stack need not be stored under all of the stacks, as it may be stored anywhere other than within the stack. Thus, the WP bit may be placed between stacks (e.g., between stack 1 and stack 2, or between stack 2 and stack 3 in FIG. 6), before all the stacks, after all the stacks, or in some other memory area entirely.
[0040] In addition, in FIG. 6, the WP management block 32 includes a WP cache 32. C Shown to include WP Cache 32 C 32 cache memory CM and Cache Control 32 CC The cache memory 32 CM is coupled to receive the stack address STKA. Consistent with the above description of the method 300, a write is desired to the stack address STKA or a read is desired from the stack address STKA. CM Also stacks 24" STKS The WP indicator bit is then combined with the corresponding WP data. Thus, the memory word of the WP indicator bit is, according to known principles, aligned with the stack space 24" STKS From cache memory 32 CM In this manner, when the stack address STKA is received, the desired WP data word (e.g., 32 bits of the WP indicator) that includes the particular WP indicator bit corresponding to STKA may already be fetched into the cache memory 32. CM , a cache hit occurs and the data is immediately available for processing (i.e., reading or writing) according to method 300. Conversely, when a stack address STKA is received, the desired WP indicator word containing the particular WP indicator bit corresponding to STKA is stored in cache memory 32. CM If it is not within the stack space, a cache miss occurs and the data is moved to the stack space. STKS From cache memory 32 CM and is then available for processing. Thus, in this example of 32-bit addressing, the least significant 5 bits of the address provide the offset for the cache access, and the most significant 27 bits provide the cache tag.
[0041] Some processors support multiple stacks simultaneously. For example, a processor may provide two stack pointers, one intended for use by the operating system and the other intended for use by applications (one example is the MSP and PSP stack pointers in some ARM processors). Another exemplary embodiment provides coverage for all concurrent stacks that the processor supports. In this embodiment, the WP bit is expanded to multiple bits represented by SPID:WP, where SPID is a set of bits large enough to uniquely encode a value for each stack pointer in the processor, and WP is a write-protect bit. During execution in the processor, the processor keeps track of which stack pointers are active. The WP management block 32 monitors signals from the processor to determine which stack pointers are in use. When the WP bit is written for an address, the SPID is also written by the WP management block 32. When the WP management block 32 has an indication from the processor that an address has been popped, the WP management block 32 clears the SPID and WP bits for that address only if the currently active stack pointer matches the SPID bits stored for that address. The WP management block 32 also provides an enable / disable signal for each stack pointer being monitored. The WP management block 32 also provides a configuration register for each stack pointer being monitored to set where in memory to store the WP bit for that stack. Changes to the enable / disable signal memory may be blocked unless the processor is in a certain mode (such as a privileged mode). The enable / disable signal, together with the configuration register, causes the operating system to stop monitoring the stack being used by a thread / application while the operating system is performing a thread switch.The operating system may disable monitoring for stack pointer A, write the location of the WP bit in memory to the configuration register for stack pointer A, change the value of stack pointer A to the appropriate value for the thread the operating system is trying to execute, and then re-enable monitoring for stack pointer A. This sequence of operations allows a thread switch (a change in the value in the stack pointer) to occur without the risk of triggering the pop detection logic of the WP management block.
[0042] From the above, the exemplary embodiments provide an improved processor with hardware support for protection against stack buffer overflows. As a result, the exemplary embodiment processor is less susceptible to coding errors or deliberate attacks that may result in undesirable or malicious stack operations. These hardware aspects also reduce time overhead compared to software-based techniques, such as avoiding source code recompilation and reducing both test and execution time overhead. The exemplary embodiments also relieve the burden on software developers, both in reducing the need for protective software and in reducing the need for deeper understanding of the architecture that would be required to address concerns that are reduced by the exemplary embodiments. Thus, the exemplary embodiments have been shown to have many advantages, and various embodiments have been provided. As yet another advantage, the exemplary embodiments contemplate various alternatives, some of which, along with others, have been described above. For example, while the exemplary embodiments have been described in the context of a 32-bit word architecture, the exemplary embodiments may also be applied to other word length architectures. As another example, although examples are shown in which a stack WP indicator bit is stored and / or cached in the stack, other associations between protective indicators including bits, flags, or multiple bits, and their possible relationships to each other (e.g., via logical AND, OR, etc.), or other conditions may be made to hardware protect each stack location in which a program counter address is stored (e.g., pushed), thereby reducing the likelihood of untimely or malicious overwrites of such addresses. As yet another example, in one example embodiment, the WP bit prevents writes to write-protected space, but in an alternative example embodiment, a write may occur and then be detected. Thus, the overwritten data is not recoverable, but a fault is generated and the overwritten data is not allowed to be used as the erroneous target address.In yet another example, example embodiments may include the ability to cause the CPU to execute instructions (e.g., while the CPU is in a privileged state) to adjust memory addresses that are being monitored for write protection as part of a stack. For example, when running a real-time operating system (RTOS), multiple stacks are used, and while example embodiments monitor changes to the CPU's stack pointer register, if the RTOS swaps stacks, the change is preferably not interpreted as a push or pop and thus a triggering condition in example embodiment methods such as that shown in FIG.
[0043] Modifications are possible in the described embodiments and other embodiments are possible within the scope of the claims.< / stka> < / stka> < / stka> < / stka> < / stka> < / stka> < / stka> < / xx>
Claims
1. 1. A processor comprising: circuitry for executing program instructions, each program instruction being stored at a location having a respective program address; circuitry for storing an indication of the program address of a program instruction to be executed; a memory stack including a plurality of stack storage spaces, each stack storage space operable to receive a program address as a return address; a number of write protection indicators corresponding to each stack storage space; A circuit element for generating a fault; Including, The circuit element generating the fault comprises: a cache memory from which values from the plurality of write protect indicators can be fetched; a detector that detects a processor read or write to the memory stack based on an instruction bus, a trace interface signal, or other hardware signals; Including, the fault generating circuitry responds to the processor write corresponding to the write of the return address by writing a return address to the stack storage space and setting a corresponding write protection indicator; the fault generating circuitry receives a processor write request to a selected stack storage space, and if a write protection indicator corresponding to that stack storage space is available in a cache memory, a cache hit occurs, the write request is executed if the write protection indicator indicates that the selected stack storage space is not write protected, and generates a fault for the processor write if the write protection indicator indicates that the selected stack storage space is write protected. Processor.
2. 2. The processor of claim 1, wherein the respective write protection indicator indicates that the selected stack storage space is write protected if the selected stack storage space stores a return address.
3. 2. The processor of claim 1, The processor, wherein the memory stack further comprises the plurality of write protect indicators.
4. 2. The processor of claim 1, each stack storage space in the plurality of stack storage spaces a first portion for storing data that may include the return address; a second portion operable as a respective write protect indicator in the plurality of write protect indicators; , a processor.
5. 5. The processor of claim 4, The second portion comprises a single bit.
6. 5. The processor of claim 4, the second portion consisting of a single least significant bit.
7. 5. The processor of claim 4, The first portion is N 4. The processor of claim 1, wherein N is selected from the set consisting of N bits, where N is an integer equal to or greater than 4.
8. 5. The processor of claim 4, the second portion consists of a single bit; The first portion is N 4. The processor of claim 1, wherein N is selected from the set consisting of N bits, where N is an integer equal to or greater than 4.
9. 2. The processor of claim 1, The processor, wherein the memory stack further comprises at least one storage space for storing the plurality of write protect indicators.
10. 10. The processor of claim 9, Each of the plurality of stack storage spaces is N 3. A processor according to claim 1, wherein N is an integer number equal to or greater than 4.
11. 2. The processor of claim 1, A processor, wherein the circuitry for storing an indication of a program address to be executed comprises a program counter.
12. 2. The processor of claim 1, The circuitry that generates the fault further inhibits overwriting the selected stack storage space with a return address.
13. 2. The processor of claim 1, 20. A processor, wherein each of said plurality of write protect indicators comprises a single bit.
14. 2. The processor of claim 1, The circuit element that generates the fault determines whether the processor read detected by the detection unit is a pop operation to read a return address, and if it is a pop operation, clears write protection using a write protection indicator corresponding to the stack storage of the return address.
15. 2. The processor of claim 1, the memory stack comprises a first memory stack in a plurality of memory stacks, each memory stack in the plurality of memory stacks comprising a respective plurality of stack storage spaces; the processor further comprising a plurality of write protect indicators for each respective memory stack of the plurality of memory stacks; a processor, wherein the generating circuitry generates a fault in response to a processor write to a selected stack storage space, for which a respective write protection indicator in a respective one of the plurality of write protection indicators indicates that the selected stack storage space is write protected.
Citation Information
Patent Citations
Stack control system and microcomputer
JP1994168145A
A method for monitoring the execution of a software program according to its specifications.
JP2001511271A
Providing extended memory protection
US20060225135A1