Hybrid mitigation of speculation-based attacks based on program behavior

By analyzing the instruction flow and selecting appropriate mitigation schemes, the speculative execution of the processor is dynamically adjusted, thus solving the side-channel attack problem caused by speculative execution in the processor and preventing information leakage without affecting performance.

CN114402324BActive Publication Date: 2025-12-26MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080064716.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-12
Filing Date
2020-06-17
Publication Date
2025-12-26
Estimated Expiration
2040-06-17

AI Technical Summary

Technical Problem

In existing processors, vulnerabilities caused by the side effects of speculative execution allow attackers to leak confidential data through side-channel attacks. Existing mitigation methods may affect processor performance or fail to effectively prevent information leakage.

Method used

By profiling the instruction flow and selecting appropriate mitigation measures, such as delays, redo, or undo mechanisms, the impact of speculative modifications on the processor cache can be prevented. The profiler dynamically selects the most suitable mitigation measure for the current workload, including the use of speculative shadow buffers and pollution matrices.

Benefits of technology

It effectively prevents information leakage caused by speculative side-channel attacks, while maintaining the processor's efficient operation in certain use cases without affecting processor performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114402324B_ABST
    Figure CN114402324B_ABST
Patent Text Reader

Abstract

Apparatuses and methods for mitigating speculative based attacks on a processor are disclosed. In one example of the disclosed technology, an apparatus includes a processor, the apparatus having a memory positioned to store profiler data to measure at least one performance criterion for an instruction stream executed by the processor, and having control logic configured to select one of a plurality of mitigation schemes to mitigate a speculative based attack on the apparatus based on the measured performance criterion. The apparatus can include a remediation unit that can prevent a speculative side effect by implementing a delay scheme, a redo scheme, or a rollback scheme, which prevents side effect data generated by a mis-speculative instruction from becoming visible to an attacker.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Attacks like Spectre and Meltdown exploit weaknesses caused by side effects of speculative execution in processors. These vulnerabilities affect billions of computers in data centers, mobile devices, laptops, and other computers. These attacks can access secrets by exploiting processor speculation and leak sensitive data by speculatively changing secrets into processor caches. This attack is very effective, breaking software-based trust abstractions like process isolation, in-process sandboxes, and even trusted hardware enclaves (e.g., Intel SGX). Thus, there is ample opportunity to improve techniques to mitigate these attacks. SUMMARY

[0002] Apparatuses and methods for mitigating speculation-based attacks in processors are disclosed. In one example of the disclosed technology, a method of operating a processor includes profiling an instruction stream for at least one performance criterion and selecting one of a plurality of mitigation schemes for a speculation-based attack based on the performance criterion. The selected mitigation scheme is selected to improve performance of the processor while implementing measures to mitigate side-channel attacks. In some examples, the plurality of mitigation schemes for cache side-channel attacks includes at least one of a delay mechanism, a redo mechanism, and a retract mechanism. As an example, based on a branch prediction performance criterion or a cache miss, one of the plurality of mitigation schemes is selected that provides the required performance based on behavior of a most recently executed instruction in the instruction stream.

[0003] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0004] The foregoing and other aspects and features of the disclosed technology are further realized and achieved by the portions of the technology as described below. BRIEF DESCRIPTION OF DRAWINGS

[0005] Figure 1 An example computing system is illustrated in which certain methods of profiling and mitigating speculation-based attacks can be performed.

[0006] Figure 2 A multi-core computing system is illustrated in which certain examples of profiling and mitigating speculation-based attacks can be performed.

[0007] Figure 3 An example of remediating a potential side-channel cache attack is illustrated, as can be performed in certain examples of the disclosed technology.

[0008] Figure 4 FIGURE illustrates an example of profiling and mitigating a speculation-based attack, as can be implemented in certain examples of the disclosed technology.

[0009] Figure 5 FIGURE illustrates an example micro-architecture using a profiler, as can be implemented in certain examples of the disclosed technology.

[0010] Figure 6 FIGURE is a chart illustrating a table of mitigation schemes that can be selected based on cache miss rates and misprediction frequencies, as can be implemented in certain examples of the disclosed technology.

[0011] Figure 7 FIGURE illustrates an example implementation of a speculation source tracking and remediation unit in a processor, as can be implemented in certain examples of the disclosed technology.

[0012] Figure 8 FIGURE illustrates an example of a speculative shadow buffer, as can be implemented in certain examples of the disclosed technology.

[0013] Figure 9 FIGURE illustrates an example of a pollution matrix, as can be implemented in certain examples of the disclosed technology.

[0014] Figure 10 FIGURE is a flowchart outlining an example method of using a profiler to select a mitigation scheme, as can be performed in certain examples of the disclosed technology.

[0015] Figure 11 FIGURE is a diagram illustrating an example computing environment in which the disclosed methods and apparatus can be implemented. DETAILED DESCRIPTION

[0016] I. General Considerations

[0017] The present disclosure is set forth in the context of representative embodiments that are not intended to be limiting in any way.

[0018] As used in this application, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. In addition, the term “includes” means “comprises.” Furthermore, the term “coupled” encompasses mechanical, electrical, magnetic, optical, as well as other practical ways of coupling or linking items together, and does not exclude the presence of intermediate elements. Moreover, as used herein, the term “and / or” means any one item or combination of items in the phrase.

[0019] The systems, methods, and devices described herein should not be construed as limiting. Rather, the present disclosure is directed to all novel and nonobvious features and aspects of the various disclosed embodiments, alone and in various combinations and subcombinations with each other. The disclosed systems, methods, and devices are not limited to any particular aspect or feature or combination of aspects and features, nor do the disclosed systems, methods, and devices require that any one or more specific advantages be present or problem to be solved. Further, any features or aspects of the disclosed embodiments can be used in various combinations and subcombinations, with each other.

[0020] Although the operations of some of the disclosed methods are described in a particular, sequential order for convenient presentation, it should be understood that unless otherwise specified, the described ordering is not mandatory. Certain implementations of the disclosed methods can be performed in an order different than presented, and / or performed concurrently. Moreover, described operations can be performed in an order different than presented, and / or performed concurrently. Additionally, for simplicity of explanation, the illustrations can not show the various ways in which the disclosed implementations can be used with other implementations. Additionally, the description sometimes uses terms like "produce", "generate", "display", "receive", "verify", "execute", "conduct", "convert", "deter", "mitigate", and "initiate" to describe the disclosed methods. These terms are high-level descriptions of the actual operations that are performed. The actual operations that correspond to these terms will vary depending on the implementation and are readily discernable by one of ordinary skill in the art.

[0021] Theoretical, scientific, or other explanations presented herein with respect to the devices or methods of the present disclosure have been provided for purposes of better understanding, and are not intended to limit the scope of the devices and methods in the appended claims to those devices and methods that function in the manner described by such theories.

[0022] Any of the disclosed methods can be implemented as computer-executable instructions, which are stored on one or more computer-readable media (e.g., computer-readable media such as one or more optical media discs, volatile memory components (such as DRAM or SRAM), or non-volatile memory components (such as hard drives)) and executed on a computer (e.g., any commercially available computer, including smart phones or other mobile devices that include computing hardware). Any computer-executable instructions for implementing the disclosed technology, as well as any data created and used during implementation of the disclosed embodiments, can be stored on one or more computer-readable media (e.g., computer-readable storage media). The computer-executable instructions can be, for example, a special purpose software application or a part of a software application that is accessed or downloaded via a web browser or other software application (such as a remote computing application). Such software can be executed, for example, on a single local computer or in a network environment (e.g., via the Internet, a wide-area network, a local-area network, a client-server network (such as a cloud computing network), or other such network) using one or more network computers.

[0023] For clarity, only certain selected aspects of the software-based implementations are described. Other aspects will be apparent from the description. For example, it should be appreciated that the disclosed technology is not limited to any particular computer language or program. For instance, the disclosed technology can be implemented by software written in C, C++, Java, or any other suitable programming language. Certain details of suitable computers and hardware are well known and need not be set forth in detail.

[0024] Furthermore, any software-based embodiments (including, for example, computer-executable instructions for causing a computer to implement any of the disclosed methods) can be up loaded, downloaded, or remotely accessed through an appropriate communication means. Such appropriate communication means include, for example, the Internet, the World Wide Web, an intranet, software applications, cable (including fiber optic cable), magnetic communication, electromagnetic communication (including RF, microwave, and infrared communications), electronic communication, or other such communication means.

[0025] II. Introduction to the Disclosed Technology

[0026] Speculative execution is used in many modern processors to avoid control flow or data dependency stalls. However, in the case of a mis-speculation, a potentially illegal access to confidential data can be temporarily allowed. For example, a side-channel attack based on latency differences of cache hits or misses can leak data to an attacker. The apparatus and methods disclosed herein can be used to address such speculative side-channel attacks by identifying the source of a speculation, monitoring the speculative execution, and remediating the side effects of the speculative execution before the speculation-source operation is resolved. Unlike mitigation approaches that attempt to prevent all speculative modifications to processor state, however, to prevent changes that can leak information in the case of a mis-speculation, the disclosed examples can use a selected approach from a number of different approaches that reduce performance in some use cases but not all.

[0027] In some examples, a number of different approaches can be used to prevent speculative modifications to processor caches such that a mis-speculation does not cause changes that can leak private information. Three such approaches include a delay approach in which a speculative load instruction is delayed until the speculation-source operation is resolved and the associated speculative load becomes non-speculative, a redo approach in which cache hits are allowed to proceed without delay but all speculative cache misses are prevented and then re-executed when the associated load instruction becomes non-speculative, and a retract approach in which speculative changes to the cache are allowed but are retracted if the speculation-source operation is determined to be a mis-speculation when the speculation-source operation is resolved. In certain examples, in addition to the retract, redo, and delay approaches, a speculative shadow buffer and / or a taint matrix can also be one of the selected mitigation approaches. In some examples, a profiler can be used to select between one or more approaches, with two different parameters used to implement at least one of the approaches. For example, one selectable delay approach can cause the processor to delay issuance or execution of a speculative load while a second selectable delay approach can cause the processor to execute but delay writeback or commit of the speculative load. Using a profiler, properties of the processor workload such as cache hit or miss rates and / or branch misprediction rates can be used to dynamically select one of the multiple approaches that is more likely to be suitable for the current workload.

[0028] As used herein, the term "speculation-source operation" refers to an operation on which speculation can be based. For example, a branch instruction introduces a control flow condition and a taint-source operation that can be performed before the speculation-source operation is resolved (e.g., whether the branch is taken or the branch location is resolved) based on whether the branch is taken. As another example, a store address computation is one example of a speculation-source operation (as used in this application) because a taint-source operation can be performed before the computation of the store address.

[0029] For ease of explanation, the examples disclosed herein focus primarily on control flow speculation for accessing confidential data by bypassing existing protection mechanisms, installing data in a cache, and subsequently leaking the data using the cache side channel. However, as will be readily appreciated by one of ordinary skill in the relevant art having the benefit of the present disclosure, the disclosed techniques can be applied to a variety of different speculation sources and side channel attacks. Examples of sources of speculation that can be addressed using the disclosed methods and apparatus include control flow speculation, data flow speculation, memory coherency, and exception checking. Examples of side channels that can be remedied from attacks based on such speculation sources can include side channel leakage involving data caches, multithreaded port attacks, translation lookaside buffer (TLB) lookups, instruction caches, use of vector instructions, and branch target buffer attacks. As used herein, the term "operation" refers not only to architecturally visible instructions (macroinstructions), but can also include processor microinstructions, microcode, or other forms of operations performed by a processor.

[0030] III. Example Computer System

[0031] Figure 1 is a block diagram 100 of an example computing system 110 in which certain examples of the disclosed technology can be implemented. The computing system 110 includes a processor having two processor cores 115 and 116, and an Ll cache 120, a shared L2 cache 125, and includes a memory and input / output unit 135. Further details are illustrated for the first core 115; the second core 116 can have similar features. As shown, the first core 115 includes control logic 130 including a branch prediction unit 140, a profiler 150, and a speculation remediation unit 160. The control logic 130 controls the operation of an execution unit 170 and a load-store unit 180 of the processor core. The execution unit 140 can include integer units, arithmetic and logic units, floating point units, vector processing units, and other appropriate data processing execution units. The load-store unit 180 can include logic circuitry that orders memory load and store operations, controls the Ll cache, and implements other control logic related to core memory operations. The control logic 130 includes an instruction scheduler that controls the dispatching and issuance of processor instructions to the execution unit 170. The control logic 130 is also coupled to the speculation tracking and remediation unit 160.

[0032] The computing system 110 and the processors (including the processor cores 115 and 116) can be implemented using any suitable computing hardware. For example, the computing system and / or processors can be implemented with a general-purpose CPU and / or a specialized processor such as a graphics processing unit (GPU) or tensor processing unit (TPU), an application-specific integrated circuit (ASIC), or programmable / reconfigurable logic such as a field-programmable gate array (FPGA) executing on any suitable commercial computer, or any suitable combination of such hardware. In some examples, the processors can be implemented as virtual processors executing on physical processors under the control of a hypervisor. In some examples, the processors can be implemented using hardware or software emulation to execute at least some instructions formatted in an instruction set architecture different from the native instruction set of the host processor providing the instruction emulation.

[0033] The control logic 130 and its subcomponents, including the branch prediction unit 140, the profiler 150, and the speculation tracking and remediation unit 160, can be implemented using any suitable techniques. The control logic 130 can be configured to regulate one or more aspects of processor control, including the execution of processor instructions through various stages of execution (e.g., fetch, decode, dispatch, issue, execution, writeback, and commit); control of the operation of the data path, execution units, and memory. The control logic 130 can regulate not only operations that are architecturally visible, but also microarchitectural operations that are not generally intended to be visible to programmers, including speculative execution (e.g., of conditional branches, memory loads, or memory address computations) from command issue, register allocation and renaming, superscalar operations, macroinstruction to microinstruction translation, fusion of macro or micro-operations, cache and memory access, branch prediction, address generation, store forwarding, instruction reordering, and any other suitable microarchitectural operations.

[0034] Control logic 130 can be implemented with "hardwired logic" (such as a finite state machine implemented with a combination of combinatorial logic gates and sequential logic gates (e.g., implemented in a random logic design style as a Moore or Mealy machine)), or as programmable logic (e.g., a programmable logic array or other reconfigurable logic); or as a microprogrammed controller or microcode processor that executes microinstructions stored in a microcode memory, which can be implemented as volatile memory (e.g., registers, static random access memory (SRAM), dynamic random access memory DRAM), non-volatile memory (e.g., read only memory (ROM), programmable read only memory (PROM), electrically erasable programmable memory (EEPROM), flash memory, etc.), or some combination of volatile and non-volatile memory types. Control logic 130 generally operates by accepting input signals (e.g., by receiving at least one digital value), processing the input signals in view of the current state of the sequential elements of the control logic, and producing output signals (e.g., by generating at least one digital value) that are used to control other components of the processor, such as logic components, data path components, execution units, memory, and / or input / output (I / O) components. Based on the input signals and the current state, the current state of the control logic is updated to a new state. Values representing the state of the control logic can be stored in any suitable storage device or memory, including latches, flip-flops, registers, register files, memory, etc. In some examples, the control logic is clocked by one or more clock signals, which allows the processing logic values to be synchronized according to clock signal edges or signal levels. In other examples, at least a portion of the control logic can operate asynchronously.

[0035] The term "conditional branch" refers to a branch that is taken or not taken based on a condition value. For example, in some instruction set architectures, another instruction is used to generate a Boolean value by comparing or testing two data (e.g., greater than, greater than or equal to, less than, less than or equal to, equal to, etc.). Depending on the Boolean value, a particular branch instruction can take the branch to a new program counter location. If the branch is not taken, the program counter is incremented (or decremented) and the next instruction in memory is executed. In some examples, a branch instruction can be predicted based on the value generated by another instruction. In some examples, an absolute branch (an instruction that specifies no condition, so will always branch when executed) can be conditional if it depends on a speculative source produced by another instruction (e.g., a memory address calculation).

[0036] The speculative tracking and remediation unit 160 acts in concert with the control logic 130 to identify sources of speculative execution in the processor, track instructions that access processor resources speculatively based on the associated source of speculative execution, and remediate the side effects of such speculative execution to reduce or eliminate the risk of side-channel attacks caused by speculative execution. In particular, the speculative tracking and remediation unit 160 can associate speculative sources with operations that cause side effects, such as memory loads, and use these associations to selectively remediate the side effects of the associated operations without having to force a delay of the entire class of operations or otherwise subjecting them to remediation measures. The speculative tracking and remediation unit 160 uses a mitigation scheme selected from a plurality of schemes, the selection being based on at least one performance metric of the instructions executed by the core 115. The speculative tracking and remediation unit 160 and its subcomponents 170, 180, and 190 can be implemented using similar hardware components as the control logic 130 as described above. In some examples, some or all of the hardware components used to implement the control logic are shared or overlap with the hardware components used to implement the speculative tracking and remediation unit 160, while in other examples, separate hardware components can be used.

[0037] In more detail, the speculative tracking and remediation unit 160 can identify and monitor one or more of many different types of operations, including, for example: control flow operations, data flow operations, branch operations, predicate operations, memory store address computations, memory coherency operations, composite atomic operations, flag control operations, transaction operations, or exception operations. Specific examples of control flow operations include branch instructions, such as relative- addressed branches and absolute-addressed jump instructions. Branches can be conditional or non-conditional (always taken or never taken). In some cases, non-conditional branches can have a speculative source, for example, when the branch instruction is pending an address computation. Even the behavior of non-conditional branches can be data-dependent, for example, in the case of branches to illegal addresses or protected locations. As another example, memory address computation operations (e.g., computations of memory addresses for memory store instructions) are another example of a speculative source that can be tracked by the speculative tracking and remediation unit 160. In some examples, a speculative shadow buffer can be used to track the source of speculation.

[0038] The speculation tracking and remediation unit 160 can identify processor operations that can be executed at least partially speculatively based on the identified speculation source. For example, memory operations such as those performed when executing memory load or memory store instructions can be executed speculatively until the speculation-source operation identified by the speculation tracking and remediation unit 160 completes. One specific example of a side-effect causing operation is a memory array read operation. Other examples of types of side-effect causing operations that can be performed prior to resolving a speculation source include: a memory load operation, a memory store operation, a memory array read operation, a memory array write operation, a memory store forward operation, a memory load forward operation, a branch instruction (including a relative addressing control flow change or an absolute addressing control flow change), an assert instruction, an implied addressing mode operation, an immediate addressing mode operation, a register addressing mode memory operation, an indirect register addressing mode operation, an auto-indexed (e.g., auto-increment or auto-decrement addressing mode operation), a direct addressing mode operation, an indirect addressing mode operation, an indexed addressing mode operation, a register-based indexed addressing mode operation, a program counter relative addressing mode operation, or a base register addressing mode operation. In some examples, a pollution matrix is used to track pollution-source operations.

[0039] The speculation tracking and remediation unit 160 is used to remediate undesirable side effects of the speculative execution. For example, accesses to the L2 cache 125 can be modified during the speculative execution so that all cache misses are blocked (a stall scheme), all cache misses are re-executed (a redo scheme), or cache misses are undone (an undo scheme). One specific example of a remedy that can occur in a stall scheme is a delay of dispatch or issue of instructions affected by the speculative execution. However, the type of remedy is not limited to a delay of dispatch or issue. For example, the remedy instruction can be delayed at another stage in the process or pipeline, e.g., earlier, at the fetch or dispatch stage, or later, at the execute, writeback, or commit stage. Examples of processor components that can be remedied by the particular speculation state change remediation unit 160 include: a data cache of the processor, an instruction cache of the processor, a register read port of the processor, a register write port of the processor, a memory load port of the processor, a memory store port of the processor, a simultaneous multithreading logic of the processor, a translation lookaside buffer of the processor, a vector processing unit of the processor, a branch target history table of the processor, or a branch target buffer of the processor.

[0040] IV. Example Computing System

[0041] Figure 2is a block diagram 200 summarizing an example computing system 201 in which certain examples of the disclosed technology can be implemented. In the illustrated computing system 201, a processor is illustrated that includes four cores 210, 211, 212, and 213. Each of the cores 211-213 can communicate with each other and with a shared logic section 220. The shared logic system 220 includes a shared level two (L2) cache 230, a memory controller 231, a main memory 235, storage 237, and input / output 238. The shared L2 cache 230 stores data accessed from the main memory 235 and can be accessed by the LI cache in each of the four cores 210-213. The memory controller 231 controls the flow of data between the shared cache 230 and the main memory 235. Additional forms of storage, such as a hard disk drive or flash memory, can be used to implement the storage 237. The input / output 238 can be used to access peripheral devices or network resources, among other appropriate input / output devices.

[0042] In Figure 2 One of the cores, core 1 210, is illustrated in more detail in FIG. 2. The other cores can have a similar or different composition than core 1 210. As shown, core 1 210 includes control logic 240 that controls operation of that particular processor core. The control logic 240 includes an instruction scheduler 241 that can control the dispatch and issue of instructions to execution units 250. The control logic 240 also includes an exception handler 242 that can be used to handle hardware or software based exceptions. The control logic 240 also includes multi-threading control logic 243 that can be used to control access to resources when the processor core is operating in a multi-threaded mode. The control logic also includes branch control logic 244 that controls the evaluation and execution of branch instructions by the processor. For example, the branch control logic 244 can be used to evaluate and execute control operations of relative branch instructions, absolute branch instructions, and / or predicate instructions. In some examples, the branch control logic 244 includes a branch history table and a branch prediction unit. The branch history table and / or branch prediction unit can be used to generate a prediction that enables speculation in the processor. For example, the branch control logic 244 can predict that a particular branch instruction will be taken or not taken, and can speculatively execute additional instructions within an instruction window based on that prediction. In some examples, the branch prediction is associated with a particular instruction in the instruction window. In other examples, the branch prediction is based on running statistics of whether branches have been taken within a particular number of instructions in the instruction window.

[0043] Execution unit 250 is configured to perform computations when executing operations, such as those specified by processor instructions. In the illustrated example, execution unit 250 includes integer execution unit 255, floating point execution unit 256, and vector execution unit 257. Integer execution unit can be used to perform integer arithmetic operations, such as addition, subtraction, multiplication, or division, shift, and rotate operations, or other suitable integer arithmetic operations. In some examples, integer execution unit 255 includes an arithmetic logic unit (ALU). Floating point execution unit 256 can perform floating point operations of single, double, or other precisions. Vector execution unit 257 can be used to perform vector operations, for example, single instruction multiple data (SIMD) instructions according to a particular vector instruction set. Examples of vector instructions include, but are not limited to, Intel SSE, SSE2, AVX, and AVX2 instruction sets; ARM Neon, SVE, and SVE2 instruction sets; PowerPC AltiVec instruction set; and certain vector examples of NVIDIA and other companies’ GPM instruction sets.

[0044] Processor core 210 also includes memory system 260, which includes level 1 (LI) instruction cache 261, LI data cache 262, and load-store unit 263. Instruction cache can be used to store instructions fetched from shared logic portion 220. Similarly, data cache can store source operands for operations performed by the processor core, and can also access memory via a memory controller in shared logic resources 220. Load-store unit 263 can regulate the operation of instruction cache 261 and data cache 262. For example, certain examples of load-store units include logic circuitry that orders memory load and store operations, controls LI instruction cache and LI data cache, and implements other control logic related to core memory operations. Load-store unit 263 uses translation lookaside buffer (TLB) 264 to translate logical addresses to physical addresses for accessing first and / or second level caches 261, 262, and / or 230. In some examples, shared logic portion 220 includes a TLB that translates logical addresses to physical addresses, instead of or in addition to the TLBs in individual cores 210-213.

[0045] Processor core also includes register file 270, which stores programmer-visible architecture registers referenced by instructions executed by the processor. Architecture registers are distinct from micro-architecture registers, and are generally specified by the instruction set architecture of the processor, while micro-architecture registers store data used when executing instructions, but are generally not programmer-visible data.

[0046] The computing system 201, including the various cores 210-213, control logic 240, memory controller 231, and other associated components, can be implemented using similar hardware components as the computing system 110, cores 115 and 116, control logic 130, and speculation tracking and remedy unit 160, as described in further detail above.

[0047] V. Example Remedies for Speculative Side Effects

[0048] Figure 3An example of source code 300 that can be compiled and executed to present a vulnerability that can be remedied by certain examples of the disclosed technology is illustrated. This source code 300 is an example of a Spectre-Vl exploit. When the code is executed, the comparison in the if statement condition of the branch instruction 310 is created as an array length check that causes the branch instruction to be executed by the processor running the code. Because the condition can take some time to evaluate, the processor can speculatively execute the instructions by predicting that the code within the braces will be executed. Thus, even if the value of the variable offset is greater than or equal to the array length, some of the instructions can be executed (but not committed) by the microarchitecture. Thus, the speculative execution of the speculative array operation 320 (arrl[offset]) will pollute the destination register of the speculative load (value). The memory value stored at the out-of-bounds address arrl[offset] can be a confidential value created by another process. This can be exploited by an attacker. For example, if the load instruction uses the confidential value as an address for a subsequent load (arr2[value] 330), subsequent accesses to that memory location will result in a cache hit. For example, after the speculative load is executed to access the confidential data, additional code can be executed 340 to iterate and attempt to load each value in the array arr2. The latency to access those values can be measured by an attacker using a timer. Thus, when the cache line hits the offset 350 corresponding to the confidential value (5), the latency of the cache will be greatly reduced. From this information, it can be inferred that the confidential value is 5. Certain processors implemented in accordance with the disclosed technology can be configured to track registers that have received data from a contaminated load. The processor can be configured to propagate the contamination when the speculative value is consumed by a subsequent instruction, thereby polluting the target registers of those instructions. Based on whether the registers are contaminated, the load is blocked in the instruction scheduler to prevent speculative cache state changes. After the original load (contaminated source) becomes non-speculative and the contamination is resolved, the load can become unblocked. Thus, in the case of a false speculation, there is no visible change in the cache based on the contaminated secret, thereby preventing information from leaking from the cache. Certain processors implemented in accordance with the disclosed technology can be configured with a remediation unit that can address this information leakage by adjusting the dispatch, issue, or execution of operations (such as loads) that cause side effects using a stall, redo, or undo mechanism.

[0049] Figure 3An example of a speculative side-effect attack is also depicted in the context of examples of how speculative tracking and remediation unit 160 can be used to mitigate attacks. A polluting-source operation can be tracked by processor's speculative tracking and remediation unit 160. For example, using speculative source tracking unit 370, speculative-source instructions such as branch instruction 310 can be monitored to determine which instructions are speculatively executed based on the speculative-source operation and to determine when the polluting-source operation condition is resolved. Similarly, speculative secret access tracking unit 380 can track polluting-source operations 315, 320, and 330 so that as the issuance and execution proceeds, the propagation of potential polluting operations through the processor can be tracked. Both units 370 and 380 can send signals to speculative state change remediation unit 390, which can determine how to resolve the potential speculative access. In some examples, a signal indicating a remediation mechanism 395 is generated and sent to remediation unit 390. For example, if it is determined that there is a speculative access to a polluted register, the remediation unit can resolve the access by adjusting the dispatch, issuance, or execution of the polluting operation using a stall, redo, or undo mechanism. The information in the profiler can be used to select this mechanism from a number of mitigation schemes, as will be discussed in further detail below. In some examples, the signal indicating remediation mechanism 395 can indicate one or more schemes, where at least one of the schemes is implemented using two different parameters. For example, one optional stall scheme can cause the processor to stall the issuance or execution of the speculative load, while a second optional stall scheme can cause the processor to execute, but stall the writeback or commit of the speculative load.

[0050] VI. Example Use of Profiler Selected Mitigation Scheme

[0051] Figure 4 FIG. 400, which illustrates processor components that can be used to mitigate side-channel speculative attacks, and illustrates three example mitigation schemes for such attacks, as can be implemented in certain examples of the disclosed technology. The components used in this mitigation include branch prediction unit 140, profiler 150, and speculative tracking and remediation unit 160. These components are used to modify the operation of illustrated load store unit 180, LI cache 120, and shared L2 cache 125 when certain operations are executed speculatively.

[0052] The three example schemes include a stall scheme 410, a redo scheme 430, and an undo scheme 450. FIG. 400 illustrates how the operation of a memory load proceeds with respect to load-store unit 180, LI cache 120, and L2 cache 125.

[0053] The illustrated delay scheme 410 shows one example of speculative mode operation that is mitigated by delaying memory loads until the operation becomes non-speculative. Thus, when there is a load, information is loaded from the Ll cache, L2 cache, or main memory, but providing the data to the load-store unit is delayed until the speculative status of the load is resolved, and it is determined that the speculative source will actually result in the memory load operation being performed. Thus, the use of speculative memory loads is delayed, and the dependent operation wake-up is delayed until the load becomes non-speculative. However, this delay scheme 410 adversely affects compute-bound workloads with more Ll hits, as additional delay is introduced in converting single-cycle Ll hits into multi-cycle operations, with branch resolution delay being padded onto the Ll hit latency. The impact is seen in Figure 4 Figure 4 The slowdown of the delay scheme is described. With the delay-based approach, workloads with high Ll hit rates are slowed down significantly.

[0054] The illustrated redo scheme 430 shows one example of speculative mode operation that is mitigated by replaying all Ll cache misses non-speculatively. As shown, the approach allows speculative Ll cache hits to proceed without delay, as they do not change the state of the Ll cache, but the scheme blocks all speculative Ll cache misses, then re-executes them after they are resolved to be non-speculative. As a result, Ll cache hits are not subject to delay, however, speculative Ll cache misses are adversely delayed, as a result, high branch resolution time is serialized with high Ll cache miss latency periods, as shown in Figure 4

[0055] The illustrated undo scheme 450 shows one example of speculative mode operation that is mitigated by allowing speculative changes to the cache, but undoing them when the miss speculation. Thus, if the workload has a high branch miss prediction rate, the undo-based approach can incur high performance overhead, as the undo mechanism can need to be invoked more frequently. Thus, if the branch miss prediction rate is relatively low, the performance of a processor using the undo mitigation scheme operation will improve.

[0056] VII. Example Use of Profiler in Speculative Remedies

[0057] Figure 5 ​​is a diagram 500 illustrating an example use of a profiler to select a speculation remediation scheme, as can be implemented in certain examples of the disclosed technology. In the illustrated example, a mitigation scheme for a processor cache is selected. In other examples, mitigation schemes for other processor components affected by speculative operations can be implemented.

[0058] As Figure 5 As shown in FIG. 1, a branch prediction unit 140 generates branch predictions for control flow instructions executed by the processor. Speculative execution occurs based on the branch predictions. However, branch predictions are not always correct, so a counter implemented in the branch prediction unit 140 can be used to count statistics of the accuracy of the branch predictions. Signals indicating whether a branch was predicted (taken or not taken), the predicted branch outcome (whether the predicted branch was taken) can be sent to the profiler 150. In other examples, other statistics, including the number of branch hits, the number of branch misses, the rate or ratio of branch hits, and the number or ratio of branch misses, are examples of statistics that can be collected during operation of the processor and sent to the profiler 150. In some examples, a program counter (PC) associated with the branch prediction can also be sent to the profiler 150. The PC can be used to associate branch prediction accuracy with a particular portion or particular instruction of the object code being executed.

[0059] One or more processor caches can also send information to the profiler 150. For example, as shown in FIG. 1, the LI cache 120 sends data to the profiler indicating cache performance. The number of cache hits, the rate of cache hits, the number of cache misses, and the rate of cache misses are examples of statistics that can be collected during operation of the processor and sent to the profiler 150. In some examples, the cache can also send data indicating one or more addresses associated with a cache hit or miss. Figure 5

[0060] ​The profiler 150 uses data from the branch prediction unit and / or the cache to generate aggregate statistics such as branch misprediction rate, LI cache hit rate, or LI cache miss rate. These aggregate statistics can be sent to the speculation tracking and remedy unit 160 to generate a mitigation policy. In some examples, the profiler 150 uses real-time data from the branch production unit and / or the cache to generate a mitigation policy that contains the real-time execution state of the processor. The mitigation policy generates a signal that indicates a selected mitigation scheme to the LI cache 125. For example, the mitigation policy signal can indicate that one of a stall scheme, a redo scheme, or a retract scheme is selected for mitigating the side effects of speculative execution. In certain examples, in addition to the retract, redo, and stall schemes, a speculative shadow buffer and / or a pollution matrix can be one of the selected mitigation schemes. In some examples, the mitigation policy signal can indicate one or more schemes, where at least one of the schemes is implemented using two different parameters. For example, one optional stall scheme can cause the processor to stall the issuance or execution of a speculative load, while a second optional stall scheme can cause the processor to execute, but stall the writeback or commit of the speculative load. The mitigation policy can also be modulated by static hints from the program regarding the security sensitivity of the memory locations being accessed. For example, static hints generated by a compiler and / or profiler data based on profiling instructions from a previous run of the program can be used to generate a default or preliminary mitigation policy, or in conjunction with real-time data collected by the hardware profiler of the processor core.

[0061] An example of a table used to select a mitigation scheme is shown in Figure 6 As shown, when the miss prediction frequency is high and the LI cache miss rate is high, a stall scheme 410 is selected. When the misprediction frequency is high, but the LI cache hit rate is relatively high (LI cache miss rate is low), then a redo scheme 430 is selected. When the misprediction frequency is relatively low, then a retract scheme 450 is selected, regardless of the cache miss rate. As will be readily understood by one of ordinary skill in the relevant art having the benefit of this disclosure, other methods or parameters can be used to select a mitigation scheme. For example, in an example where only two mitigation schemes are available, a different table than the one shown in Figure 6 will be used to select a mitigation scheme. In other examples, there can be more mitigation schemes available, so the example table will be modified as appropriate. In some examples, the values used in the illustrated table are predetermined and programmed into the processor itself. In other examples, the mitigation scheme table can be dynamically adjusted based on feedback from a programmer, an operating system, or other appropriate method that selects a more appropriate mitigation scheme based on the current or past workload of the processor.

[0062] VIII.Example Pollution Matrix Mitigation Scheme

[0063] Reference is made below Figures 7-10 A further detailed example of a particular version of a delay-based mitigation scheme using a pollution matrix is discussed. In particular, this particular delay-based mitigation scheme uses speculative shadow registers and a pollution matrix in order to delay some, but not all, memory operations in order to address speculative side-channel attacks. Those of ordinary skill in the art having benefit of the present disclosure will readily appreciate that other delay schemes can be implemented that do not use the fine-grained pollution matrix discussed in this section. For example, a simpler delay scheme can delay all operations that can be polluted until the speculation source is resolved without using fine-grained tracking using speculative shadow buffers and a pollution matrix.

[0064] Figure 7 is a block diagram 700 summarizing an example processor microarchitecture in which certain examples of the disclosed technology can be implemented. As Figure 7 As shown in Figure 7 , the control logic 710 also includes a branch predictor 750. The branch predictor can monitor speculation-source operations performed by the processor to predict whether branch instructions, and other appropriate speculation-source instructions, will execute, or whether their branches will be taken.

[0065] Also in Figure 7An example set of hardware for performing processor operations is shown in FIG. 8. This includes an instruction fetch unit 760 and an instruction decoder 762, as well as a dispatch and issue unit 763. The instruction fetch unit 760 is used to fetch instructions from memory or an instruction cache. The instruction decoder 762 decodes fetched instructions and generates control signals for configuring the processor and gathering input operands for processor operations. The dispatch and issue unit 763 dispatches particular operations to particular execution units 770 of the processor. The dispatch and issue unit 763 also controls when instructions are allowed to issue for execution by the execution units 770. Also shown in FIG. 8 are a register file 780 and a load store queue 785. The processor also includes a memory subsystem, including an LI cache 790, an L2 cache 792, and a memory 795. Figure 7 The register file 780 and load store queue 785 are shown in FIG. 8. The processor also includes a memory subsystem, including an LI cache 790, an L2 cache 792, and a memory 795.

[0066] The control logic 710, including the speculation source tracking and remedy unit 720, the issue suppressor 737, the dynamic instruction scheduler 740, the branch predictor 750, and other associated components, can be implemented using similar hardware components as the computing system 110, the cores 115 and 116, the control logic 130, and the speculation tracking and remedy unit 160, as described in further detail above.

[0067] Figure 8 is a block diagram 800 outlining aspects of an example speculation source tracking unit, as can be implemented in certain examples of the disclosed technology. As Figure 8 As shown in FIG. 8, a reorder buffer (ROB) stores tags that indicate the number of processor instructions that have been ordered to be executed as shown from right to left. For example, a first load instruction LI is issued first, followed by a conditional branch instruction Bl, a second load instruction L2, a store instruction SI, and a third load instruction LI.

[0068] The speculative source tracking unit includes a speculative shadow buffer 820. The speculative shadow buffer 820 stores indicators of instructions in the ROB 810 that have been identified as speculative sources. Thus, the branch instruction is stored at the head of the speculative shadow buffer 820, followed by the store instruction S1. As described above, the branch instruction B1 pollutes all instructions in the ROB 810 following it until its associated speculation-source operation (which determines whether the branch will be taken, or in some cases determines the address of the target branch) has been resolved, and thus the subsequent instructions are no longer considered speculative. Similarly, the store instruction S1 pollutes all instructions in the ROB 810 following it until its associated speculation-source operation has been resolved, e.g., the computation of the address to which data for execution of the store instruction S1 will be stored, the gate resolving the instruction, and any instructions that depend on the store instruction S1. In addition, instructions in the load queue 830 can be associated with speculative sources. In the illustrated example, the second load instruction L2 is identified as speculative because until the speculation-source operation associated with the branch instruction B1 is resolved, it is not known whether the instruction will execute and commit. Similarly, the third load instruction L3 is speculative until the previous polluting-source operations S1 and B1 are resolved. When the associated speculation-source instruction executes and commits, the entry can be removed from the speculative shadow buffer 820, and the remediation unit can take appropriate action to complete mitigation of the operation that caused the side effect. For example, if a delay mitigation is selected, appropriate load data can be forwarded to dependent instructions. Otherwise, if a redo mitigation is selected, the load can be safely replayed when the speculation-source instruction has executed. If a retract mitigation is selected, there is no longer a need to retract the side effect of the load.

[0069] Figure 9is a diagram 900 illustrating an example pollution matrix 910 that can be used to track sources of speculative pollution, in accordance with certain examples of the disclosed technology. In the illustrated example, each column in the pollution matrix 910 is associated with an architectural register of a processor (e.g., Rl, R2, R3, etc.). Each row of the pollution matrix 910 is associated with a load instruction (e.g., LI, L2, L3, etc.). In typical implementations, where a processor microarchitecture implements register renaming, the register columns are associated with logical processor registers or physical processor registers. For memory load operations, a tag or other identifier can be used to track which particular load instructions are associated with particular columns of the pollution matrix 910. As illustrated, the pollution matrix stores associations between memory load instructions and registers affected by the load instructions. For example, the first row indicates that load instruction LI is associated with register R2. This is a typical case where a memory load instruction writes its result to register R2. The second row indicates that load instruction L2 is associated with polluting register R3. The third row indicates that a single load instruction L3 has a pollution tag associated with two registers, Rl and R2. This is because, as will be discussed further below, subsequent instructions that use values that can be polluted can also be tagged as being polluted. Thus, when a speculation source is resolved, one or more registers that are tracked as being polluted can not be polluted as part of a remediation process.

[0070] IX. Example Method of Mitigating Side Effects Using Mitigation Scheme Selected According to Performance Criteria

[0071] Figure 10 is a flowchart 1000 summarizing an example method of mitigating side effects using a mitigation scheme selected according to performance criteria, as can be implemented in certain examples of the disclosed technology. For example, a processor including a speculation tracking and remediation unit, such as those discussed above, can be used to perform the illustrated method.

[0072] At process block 1010, the instruction stream is profiled for at least one performance criterion. For example, statistics related to control flow, such as branch misprediction rates, and statistics related to performance of memory structures, such as caches, including cache hit or cache miss rates, can be collected by the profiler. Generally, the performance criterion will vary for a particular instruction stream based on the amount of speculative execution that occurs. Thus, some target code can exhibit higher or lower branch misprediction and / or cache hit or miss rates. In some examples, the profiling is performed dynamically during runtime operation of the processor. In some examples, hardware, such as hardware performance counters or past behavior counters, can be used in conjunction with the statistics of the profiler. In some examples, the at least one performance criterion relates to branch prediction, and the profiling is performed using saturating counters, Lee-Smith counters, pattern history tables, branch history tables, or global history tables with index sharing. In some examples, the performance criterion is based on accuracy of branch prediction.

[0073] At process block 1020, based on the performance criteria collected at process block 1010, one of a plurality of mitigation schemes is selected for use in mitigating the speculative-based attack. In some examples, the selection of the mitigation is performed dynamically during runtime operation of the processor. In some examples, the mitigation scheme is selected from a plurality including a delay mechanism, a redo mechanism, and / or a rollback mechanism. In some examples, the selection is performed by measuring the at least one performance criterion while operating the processor using a first one of the mitigation schemes, measuring the at least one performance criterion while operating the processor using a second, different one of the mitigation schemes, and comparing the measurement of the at least one performance criterion while operating the processor using the first one of the mitigation schemes to the measurement of the at least one performance criterion while operating the processor using the second one of the mitigation schemes. In some examples, a table, such as the table shown in FIG. 6, is used to determine the mitigation scheme based on the relative branch misprediction rates and cache hit / miss rates. In some examples, the selection is based at least in part on a compiler hint inserted into the target code of the instruction stream, the compiler hint indicating the performance criterion. For example, the compiler hint can indicate that a plurality of portions of the code, for which the compiler has determined that one of a plurality of mitigation schemes will perform better, will be executed. In some examples, the compiler hint can indicate the performance criterion used to select one of the plurality of mitigation schemes. Figure 6

[0074] ​At process block 1030, the selected mitigation scheme is used to mitigate the side effects of the speculatively executed processor operation. For example, the mitigation can include at least one of: suppressing fetch of the speculative operation; suppressing decode of the speculative operation; suppressing dispatch of the speculative operation; suppressing issue of the speculative operation; suppressing execution of the speculative operation; suppressing memory access of the speculative operation; suppressing register write back of the speculative operation; or suppressing commit of the speculative operation. In some examples, the side effects affect the state of at least one of: a data cache of the processor, an instruction cache of the processor, a register read port of the processor, a register write port of the processor, a memory load port of the processor, a memory store port of the processor, a simultaneous multithreading logic of the processor, a translation lookaside buffer of the processor, a vector processing unit of the processor, a branch target history table of the processor, or a branch target buffer of the processor.

[0075] Examples of speculative operations that can be mitigated using the selected mitigation scheme include: a memory load operation, a memory store operation, a memory array read operation, a memory array write operation, a memory store forward operation, a memory load forward operation, a relative branch instruction, an absolute addressing branch instruction, an assert instruction, an implied addressing mode operation, an immediate addressing mode operation, a register addressing mode memory operation, an indirect register addressing mode operation, an auto-indexed addressing mode operation (including addresses computed by incrementing or decrementing a base address), a direct addressing mode operation, an indirect addressing mode operation, an indexed addressing mode operation, a register-based indexed addressing mode operation, a program counter relative addressing mode operation, or a base register addressing mode operation. Further, the speculation can be based on many different speculation sources, including operations that are executed speculatively based on conditional operations, including at least one of: a control flow operation, a data flow operation, a branch operation, an assert operation, a memory store address computation, a memory coherency operation, a composite atomic operation, a flag control operation, a transaction operation, or an exception operation.

[0076] X. Example General Purpose Computing Environment

[0077] Figure 11 FIG. 1 illustrates a generalized example of a computing environment 1100 in which the described embodiments, methods, and techniques, including selecting a mitigation scheme and using a mitigation scheme selected based on processor performance criteria to mitigate speculative operation side effects, can be implemented.

[0078] The computing environment 1100 is not intended to suggest any limitation as to scope or functionality of the technology, as the technology can be implemented in diverse general- purpose or special-purpose computing environments. For example, the disclosed technology can be implemented with other computer system configurations, including hand held devices, multiprocessor systems, micro-processor based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, etc. The disclosed technology can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote memory storage devices.

[0079] With reference to Figure 11 , the computing environment 1100 includes at least one processing unit 1110 and memory 1120. In Figure 11 , the most basic configuration 1130 is included within a dashed line. The processing unit 1110 executes computer-executable instructions and can be a real or virtual processor. In a multi-processing system, multiple processing units execute computer-executable instructions to increase processing power and thus, the multiple processors can be running simultaneously. The memory 1120 can be volatile memory (e.g., registers, cache, RAM), non-volatile memory (e.g., ROM, EEPROM, flash memory, etc.), or some combination of the two. The memory 1120 stores software 1180 that can, for example, implement the technology described herein, images, and videos. A computing environment can have additional features. For example, the computing environment 1100 includes storage 1140, one or more input devices 1150, one or more output devices 1160, and one or more communication connections 1170. An interconnection mechanism (not shown) such as a bus, controller, or network interconnects the components of the computing environment 1100. Typically, operating system software (not shown) provides an operating environment for other software executing in the computing environment 1100, and coordinates activities of the components of the computing environment 1100.

[0080] The storage 1140 can be removable or non-removable, and includes magnetic disks, magnetic tapes or cassettes, CD-ROMs, CD-RWs, DVDs, or any other medium that can be used to store information and that can be accessed by the computing environment 1100. The storage 1140 stores instructions for the software 1180, which can be used to implement the technology described herein.

[0081] Input device(s) 1150 can be a touch input device such as a keyboard, keypad, mouse, touch screen display, pen, or a trackball, a voice input device, a scanning device, or another device that provides input to computing environment 1100. For audio, input device(s) 1150 can be a sound card or similar device that accepts audio input in analog or digital form, or a CD-ROM reader that provides audio samples to computing environment 1100. Output device(s) 1160 can be a display, printer, speaker, CD-writer, or another device that provides output from computing environment 1100.

[0082] Communication connection(s) 1170 enable communication over a communication medium to another computing entity. The communication medium conveys information such as computer-executable instructions, compressed graphics information, video, or other data in a modulated data signal. Communication connection(s) 1170 are not limited to wired connections (e.g., megabit or gigabit Ethernet, wireless broadband techniques, optical fiber connections on optical fiber channels, or the like) and include wireless technologies (e.g., RF connections via Bluetooth, WiFi (IEEE 802.11a / b / n), WiMax, cellular, satellite, laser, infrared, or the like), and other appropriate communication connections for providing a network connection to software and hardware. In a virtual host environment, the communication connection(s) can be a virtualized network connection provided by the virtual host.

[0083] Some embodiments of the disclosed methods can be performed using computer-executable instructions implemented in computing cloud 1190 to implement all or part of the disclosed technology. For example, the disclosed methods can be performed on processing units located in computing environment 1130, or the disclosed methods can be performed on servers located in computing cloud 1190.

[0084] Computer-readable media are any available media that can be accessed within computing environment 1100. By way of example, and not limitation, with computing environment 1100, computer-readable media include memory 1120 and / or storage 1140. It should be appreciated that the term "computer-readable storage media" includes media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. It should be appreciated that the term "computer-readable storage media" includes media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data.

[0085] XI. Additional Examples of the Disclosed Technology

[0086] A system one or more computers can be configured to perform a particular operation or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes the system to perform the action. One or more computer programs can be configured to perform a particular operation or actions by virtue of including instructions that, when executed by the processor or other data processing apparatus, cause the apparatus to perform the action. A general aspect includes profiling an instruction stream for at least one performance criterion. The method also includes selecting one of a plurality of mitigation schemes for a speculation-based attack based on the performance criterion. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0087] Implementations can include one or more of the following features. A method in which at least one performance criterion varies due to speculative execution. In which the plurality of mitigation schemes includes at least one of: a stalling mechanism, a redoing mechanism, or a revoking mechanism. In some examples, the plurality of mitigation schemes includes at least a stalling mechanism and a redoing mechanism. In some examples, the plurality of mitigation schemes includes at least a stalling mechanism and a revoking mechanism. In some examples, the plurality of mitigation schemes includes at least a redoing mechanism and a revoking mechanism. In some examples, the plurality of mitigation schemes includes at least one of: a stalling mechanism, a redoing mechanism, or a revoking mechanism, and includes a scheme that does not mitigate speculation-based attacks, or a limiting scheme that limits all instructions that can be tainted by a speculation source. In some examples, selective mitigation using a taint matrix is employed. In some examples, the mitigation scheme is selected to be used only with certain code, threads, processes, or processor cores, while other code does not use mitigation for speculation-based attacks. For example, code, threads, processes, and / or cores that are designated as more sensitive or having a higher protection level can use a profile to select a scheme, while other aspects have a lower protection level, use a different scheme, or do not use a mitigation scheme.

[0088] Implementations can also include one or more of the following features. For the method, wherein the profiling and the selecting are performed dynamically during runtime operation of the processor. For the method, wherein the profiling is performed using a hardware performance counter of the processor. For the method, wherein the profiling is performed in real-time during execution of the program. For the method, wherein the selecting combines hints to be generated by a compiler or profiler data collected from a previous execution of the program with real-time data measured during execution of the program. For the method, wherein the at least one performance criterion relates to branch prediction, and the profiling is performed using saturating counters, Lee-Smith counters, pattern history tables, branch history tables, or global history tables with index sharing. For the method, wherein the at least one performance criterion is measured using past behavior counters. For the method, wherein: the at least one performance criterion is based on accuracy of branch prediction. For the method, wherein the at least one performance criterion is based on at least one of: cache hit rate or cache miss rate for a cache of the processor. For the method, wherein the selecting is performed by: measuring the at least one performance criterion while a first of the mitigation schemes is in use while operating the processor. The method can also include: measuring the at least one performance criterion while a second, different, of the mitigation schemes is in use while operating the processor. The method can also include: comparing the measurements of the at least one performance criterion while the first of the mitigation schemes is in use while operating the processor, with the measurements of the at least one performance criterion while the second of the mitigation schemes is in use while operating the processor. The method can also include selecting the scheme using a compiler hint inserted in the object code, the compiler hint indicating the performance criterion. The method can also include: mitigating side effects of the at least one instruction in the stream of speculatively executed instructions using the selected mitigation scheme. For the method, wherein the mitigating includes at least one of: suppressing fetch of the speculatively operated on; suppressing decode of the speculatively operated on; suppressing dispatch of the speculatively operated on; suppressing issue of the speculatively operated on; suppressing execution of the speculatively operated on; suppressing memory access of the speculatively operated on; suppressing register write back of the speculatively operated on; or suppressing commit of the speculatively operated on. For the method, wherein the at least one instruction is speculatively executed based on a conditional operation, the conditional operation including at least one of: a control flow operation, a data flow operation, a branch operation, an assert operation, a memory store address calculation, a memory operation, a composite atomic operation, a flag control operation, a transaction operation, or an exception operation.For the method, where executing the at least one instruction includes speculatively executing at least one of: a memory load operation, a memory store operation, a memory array read operation, a memory array write operation, a memory store-forward operation, a memory load-forward operation, a relative branch instruction, an absolute addressing branch instruction, a predicate instruction, an implied addressing mode operation, an immediate addressing mode operation, a register addressing mode memory operation, an indirect register addressing mode operation, an auto-indexed addressing mode operation (including addresses computed by incrementing or decrementing a base address), a direct addressing mode operation, an indirect addressing mode operation, an indexed addressing mode operation, a register-based indexed addressing mode operation, a program counter relative addressing mode operation, or a base register addressing mode operation. For the method, where the side effect affects a state of at least one of: a data cache of the processor, an instruction cache of the processor, a register read port of the processor, a register write port of the processor, a memory load port of the processor, a memory store port of the processor, a simultaneous multithreading logic of the processor, a translation lookaside buffer of the processor, a vector processing unit of the processor, a branch target history table of the processor, or a branch target buffer of the processor. Implementations of the described technology can include hardware, a method or process, or computer- executable instructions stored on a computer- accessible medium.

[0089] One general aspect includes a computer-readable storage medium storing computer-readable instructions that, when executed by a computer, cause the computer to generate a design file for a circuit that, when fabricated, causes a processor to perform a method. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, all of which are configured to perform the actions of the methods.

[0090] One general aspect includes an apparatus to implement a processor, the apparatus comprising: a memory positioned to store profiler data to measure at least one performance criterion for a stream of instructions executed by the processor; and control logic configured to:. The apparatus further comprises: select one of a plurality of mitigation schemes to mitigate a speculation-based attack on the apparatus based on the measured performance criterion. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, all of which are configured to perform the actions of the methods.

[0091] Implementations can include one or more of the following features. For the apparatus, wherein the at least one performance criterion varies due to the speculative execution. For the apparatus, wherein the plurality of mitigation schemes includes at least one of: a stall mechanism, a redo mechanism, or a void mechanism. For the apparatus, wherein the profiling and selection are performed dynamically during runtime operation of the processor. For the apparatus, wherein the profiling is performed using a hardware performance counter of the processor. For the apparatus, wherein the apparatus further includes branch prediction hardware, and wherein the at least one performance criterion relates to branch prediction, and the profiling is performed using at least one of the following branch prediction hardware: a saturating counter, a Lee-Smyth counter, a pattern history table, a branch history table, or a globally shared history table with indexed sharing. The apparatus further includes a past behavior counter, wherein the at least one performance criterion is measured using the past behavior counter. For the apparatus, wherein the processor includes a cache, and wherein the at least one performance criterion is based on a cache hit rate or a cache miss rate for the cache. For the apparatus, wherein the apparatus is further configured to perform at least one of any of the methods. For the apparatus, wherein the control logic includes a pollution matrix, and wherein at least one of the mitigation schemes uses the pollution matrix to determine dependencies to mitigate a speculation-based attack. For the apparatus, wherein the pollution matrix stores data indicating operations that depend on an identified speculative operation, and at least one of the mitigation schemes includes refraining from at least one side effect of the identified speculative operation until a committed condition state of the speculative operation is resolved. For the apparatus, wherein the control logic further includes: circuitry configured to clear pollution data in a memory to indicate whether the identified speculative operation has been resolved. The apparatus can further include an execution unit that executes the speculative operation that causes the at least one side effect. The apparatus can further include the execution unit that executes the dependent operation based on the cleared pollution data. The apparatus can further include the control logic including the pollution matrix storing data indicating at least two operations that depend on an identified speculative operation. For the apparatus, wherein the control logic includes a pollution matrix storing data indicating that the speculative operation is caused by executing a store memory instruction, and the condition state is determined by executing a branch instruction. Implementations of the described technology can include hardware, a method or process, or computer software on a computer-accessible medium. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, all of which are configured to perform the actions of the methods.

[0092] In view of the many possible embodiments to which the principles of the disclosed subject matter can be applied, it should be recognized that the illustrated embodiments are only preferred examples and should not be taken as limiting the scope of the claims. Rather, the scope of the disclosed subject matter is defined by the following claims. Therefore, all applications falling within the scope of these claims are intended to be covered.

Claims

1. A method of operating a processor, the method comprising: profiling one or more performance criteria for an instruction stream; and using a hint inserted by a compiler into object code to direct selection of one of a plurality of mitigation schemes at runtime for an attack based on speculation, wherein the hint in the object code indicates a performance criterion of the performance criteria to be used for the selection, and the selection is based on the indicated performance criterion, wherein the selection is performed by: measuring the one or more performance criteria while a first of the mitigation schemes is used in operating the processor; measuring the one or more performance criteria while a different second of the mitigation schemes is used in operating the processor; and comparing the measurements of the one or more performance criteria while a first of the mitigation schemes is used in operating the processor, with the measurements of the one or more performance criteria while a second of the mitigation schemes is used in operating the processor.

2. The method of claim 1, wherein: the one or more performance criteria vary due to speculative execution.

3. The method of claim 1, wherein the plurality of mitigation schemes includes at least one of: a redo mechanism or a rollback mechanism.

4. The method of claim 1, wherein the selection action combines a hint generated by a compiler with real-time data measured during execution of a program.

5. The method of claim 1, wherein: the one or more performance criteria are based on at least one of: accuracy of branch prediction, cache hit rate for a cache of the processor, or cache miss rate for a cache of the processor.

6. The method of claim 1, further comprising: using the selected mitigation scheme to mitigate side effects of speculative execution of an instruction in the instruction stream, the mitigating including at least one of: suppressing fetch of the instruction; suppressing decode of the instruction; suppressing dispatch of the instruction; suppressing issue of the instruction; suppressing execution of the instruction; suppressing memory access of the instruction; suppressing register writeback of the instruction; or suppressing commit of the instruction.

7. The method of claim 1, wherein the indicated performance criterion used by the selection of the selected mitigation scheme includes branch misprediction rate.

8. A computer readable storage medium storing computer readable instructions which, when executed by a computer, cause the computer to generate a design file for a circuit which, when fabricated using the design file, causes a processor to perform the method of claim 1.

9. An apparatus to implement a processor, the apparatus comprising: a memory positioned to store profiler data to measure branch misprediction rate for an instruction stream executed by the processor; branch prediction hardware; and control logic configured to: ​ ​ profile the instruction stream using the branch prediction hardware to obtain the profiler data, and selecting, based at least in part on the measured branch misprediction rate, one of a plurality of mitigation schemes to mitigate a potential speculation-based attack on the apparatus by: measuring the one or more performance criteria while a first one of the mitigation schemes is in use when operating the processor; measuring the one or more performance criteria while a different second one of the mitigation schemes is in use when operating the processor; and comparing the measurements of the one or more performance criteria when the first one of the mitigation schemes is in use when operating the processor, with the measurements of the one or more performance criteria when the second one of the mitigation schemes is in use when operating the processor.

10. The apparatus of claim 9, wherein the plurality of mitigation schemes includes at least one of: a stall mechanism, a redo mechanism, or a rollback mechanism.

11. The apparatus of claim 9, wherein: the profiling and the selecting are performed dynamically during runtime operation of the processor, at least one of the profiling or the selecting being performed using a hardware performance counter of the processor.

12. The apparatus of claim 9, further comprising a past behavior counter, wherein the branch misprediction rate is measured using the past behavior counter.

13. The apparatus of claim 9, wherein the processor includes a cache, and wherein the selecting is further based on a cache hit rate or a cache miss rate for the cache.

14. The apparatus of claim 9, wherein the control logic includes a pollution matrix, and wherein at least one of the mitigation schemes uses the pollution matrix to determine dependencies to mitigate the speculation-based attack.

15. The apparatus of claim 9, wherein the control logic further comprises: circuitry configured to clear pollution data in the memory to indicate whether an identified speculative operation has been resolved; an execution unit that executes the speculative operation that causes at least one side effect; and based on the cleared pollution data, an execution unit that executes another operation that depends on the speculative operation.

16. The apparatus of claim 9, wherein the control logic includes a pollution matrix that indicates a conditional instruction that resolves a condition state of a speculation-source operation when the conditional instruction is executed, the pollution matrix storing data that indicates at least two operations that depend on an identified speculative operation, the pollution matrix further storing data that indicates that the speculation operation was caused by executing a memory load instruction, and the condition state was determined by executing a branch instruction.

17. An apparatus comprising: means for profiling an instruction stream to be executed by a processor for at least one performance criterion; and ​ A component for mitigating a speculative-based attack by selecting one of a plurality of mitigation schemes based on the at least one performance criterion, wherein the selecting is performed by: measuring the one or more performance criteria while the first one of the mitigation schemes is in use while operating the processor; measuring the one or more performance criteria while a different second one of the mitigation schemes is in use while operating the processor; and comparing the measurements of the one or more performance criteria while the first one of the mitigation schemes is in use while operating the processor, with the measurements of the one or more performance criteria while the second one of the mitigation schemes is in use while operating the processor; A component for remediating information leakage by undoing cache changes when an error is predicted; and A component for stalling a load instruction when a speculative cache miss occurs, and re-executing the load instruction after the speculative cache miss becomes non-speculative. ​

Citation Information

Patent Citations

  • Method and structure for monitoring pollution and prefetches due to speculative accesses

    US20050251629A1

  • Using a conversion look aside buffer to implement an instruction set agnostic runtime architecture

    US20160026487A1

  • Speculative side-channel attack mitigations

    US20190114422A1

  • Methods and apparatus to detect side-channel attacks

    US20190147162A1