Processor memory reordering hints in bit-accurate tracing

By recording memory reordering prompts and processor status in the processor, debugging difficulties caused by out-of-order execution of modern processors are solved, and debugging performance and resource utilization efficiency are improved.

CN113168367BActive Publication Date: 2025-05-16MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980081491.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-10-23
Filing Date
2019-10-11
Publication Date
2025-05-16
Estimated Expiration
2039-10-11

AI Technical Summary

Technical Problem

The out-of-order execution and speculative execution of modern processors lead to cache flow into out-of-order records, which in turn increases the difficulty of determining the correct memory value during playback in time travel debugging, consuming a large amount of processor and memory resources.

Method used

Additional memory reordering prompts and processor status are recorded in the processor, providing information for identifying the correct memory value during tracking playback.

Benefits of technology

Reduces the processing time and resources required to determine the correct memory value during tracking post-processing and playback, improving the performance and utility of time travel tracking and debugging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113168367B_ABST
    Figure CN113168367B_ABST
Patent Text Reader

Abstract

Storing a memory reorder hint to a processor trace includes: while the system is executing a plurality of machine code instructions, the system initiating execution of a particular machine code instruction that performs a load of a memory address. Based on the initiation of the instruction, the system initiating storage of a particular cache line in a processor cache that stores a first value corresponding to the memory address to the processor trace. After initiating the storage of the particular cache line and before committing the particular machine code instruction, the system detects an event that affects the particular cache line. Based on the detection, the system initiates storage of the memory reorder hint to the processor trace.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] When writing code during the development of a software application, developers usually spend a lot of time "debugging" the code to find runtime errors in the code. In this case, developers can use several methods to reproduce and locate errors in the source code, such as observing the behavior of the program based on different inputs, inserting debugging code (e.g., printing variable values, tracing execution branches, etc.), temporarily removing code sections, etc. Tracking runtime errors to find out code failures may take up time in application development.

[0002] Many types of debugging applications ("debuggers") have been developed to assist developers in the code debugging process. These development tools provide developers with the ability to track, visualize, and change the execution of computer code. For example, a debugger can visualize the execution of code instructions, can present variable values ​​at different times during code execution, can enable developers to change the code execution path, and / or can enable developers to set "breakpoints" and / or "watchpoints" on code elements of interest (which, when reached during execution, will cause the execution of the code to be suspended), etc.

[0003] An emerging form of debugging application implements "time travel," "reverse," or "historical" debugging, in which the execution of one or more of a program's threads is recorded or traced by tracing software and / or hardware into one or more trace files. Using some tracing techniques, these (multiple) trace files contain "bit accurate" traces of the execution of each traced thread, which can be used to later replay the execution of each traced thread for forward and backward analysis. Using bit accurate traces, the previous execution of each traced thread can be reproduced to the granularity of its individual machine code instructions. Using these bit accurate traces, a time travel debugger enables the developer to set forward breakpoints (such as a conventional debugger), and reverse breakpoints during the replay of the traced threads.

[0004] One form of hardware-based trace recording records a bit-accurate trace based in part on recording cache hits (e.g., cache misses) to the microprocessor during execution of the machine code instructions of each traced thread by the processor. These recorded cache hits enable a time-travel debugger to later replay any memory values ​​read by these machine code instructions during replay of the traced thread.

[0005] Modern processors are generally not sequentially consistent in their memory accesses in order to ensure that the processor can stay actually busy. As a result, modern processors can reorder memory accesses based on the order in which the memory accesses appear in the machine code instruction stream. One way that modern processors may reorder memory accesses is by executing the machine code instructions of a thread out of order (e.g., in an order different from the order in which the instructions are specified in the thread's code). For example, a processor may execute multiple non-dependent memory loads and / or stores simultaneously across parallel execution units, rather than one by one as they appear in the thread instructions. Another way that modern processors may reorder memory accesses is by engaging in "speculative" execution of thread instructions before determining the condition(s) for which the outcome of a branch is actually known, such as speculatively prefetching instructions and executing the instructions after the branch. The out-of-order and / or speculative execution of thread instructions means that the memory values ​​that these instructions depend on may appear in the processor's cache from time to time, rather than being committed (and therefore recorded) from an architectural perspective when the memory access instruction is made. Furthermore, the act of speculatively prefetching instructions changes the contents of the processor's caches, even if those instructions are not actually executed, and even if they do not access memory. The extent to which a given processor participates in out-of-order and / or speculative execution may depend on the instruction set architecture and implementation of the processor. Summary of the invention

[0006] When recording the execution of one or more threads that are out of order and / or speculatively executed at a modern processor, cache inflows may be recorded out of order according to the order of the thread instructions. Due to speculative execution, these cache inflows may not even actually need to be used for the correct replay of the execution of the traced thread. Therefore, the debugger that replays the trace recorded in those processors may need to track the cache values ​​of multiple potential log records, which may be used by a given instruction and determine which presents the correct execution result. When a single correct result can be determined mathematically (for example, by solving a graph problem based on the knowledge of the future program state such as memory access, register values, etc.), the process of actually identifying this single correct result will consume significant processing time during trace post-processing and / or replay-this will reduce post-processing and / or replay performance in consuming additional processor and memory resources.

[0007] At least some embodiments described herein include modifications to a microprocessor (processor) that cause the processor that records the execution of a thread into a trace to also record additional memory reorder hints into the trace. These hints provide information that can be used during trace replay to help identify which memory value was actually used by a given memory access machine code instruction. Examples of memory reorder hints include information that helps identify how long ago a particular machine code instruction read from a particular cache line (e.g., in terms of the number of processor cycles, the number of instructions, etc.), whether a particular machine code instruction reads a current value from a particular cache line, an indication of which value a particular machine code instruction read from a particular cache, and the like.

[0008] Additionally, at least some embodiments described herein may also include processor modifications that cause the processor to log additional processor state into the trace. Such additional processor state includes, for example, the value of at least one register, a hash of at least one register, an instruction count, at least a portion of a processor branch trace, etc. Such processor state may provide additional bounds for the mathematical problem of determining which of a plurality of logged cache values ​​presents the correct execution result.

[0009] It should be understood that the embodiments described herein can reduce (or even eliminate) the processing required during trace post-processing and / or trace replay for identifying which cache values ​​of specific log records are consumed by trace thread instructions. Therefore, the embodiments described herein solve technical problems that uniquely arise in the field of time travel tracing and debugging / replay. The technical solutions described herein improve the performance of trace post-processing and / or trace replay, greatly improve the practicality of time travel tracing and debugging / replay, and reduce the processing / memory resources required during trace post-processing and / or trace replay.

[0010] In some embodiments, a system (such as a microprocessor) stores a memory reorder hint in a processor trace. The system includes one or more processing units (e.g., cores), and a processor cache including a plurality of cache lines. The system is configured to execute a plurality of machine code instructions at the one or more processing units. During execution, the system initiates execution of a particular machine code instruction, the particular machine code instruction performing a load of a memory address. Based on the initiation of the particular machine code instruction, the system initiates logging of a particular cache line in a processor cache that overlaps a memory address into the processor trace, including initiating logging of a value corresponding to the memory address associated with logging the particular cache line. After initiating logging of the particular cache line into the processor trace, and before committing the particular machine code instruction, the system detects an event that affects the cache line. Based at least on detecting the event that affects the particular cache line, the system initiates storage of the memory reorder hint to the processor trace.

[0011] This Brief Description provides a selection of concepts in a simplified form that are further described in the Detailed Description below. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to describe the manner in which the above and other advantages and features of the present invention can be obtained, a more particular description of the invention briefly described above will be presented by reference to specific embodiments of the invention shown in the accompanying drawings. It should be understood that these drawings depict only typical embodiments of the invention and are not, therefore, to be considered limiting of the scope of the invention. The invention will be described and explained with additional specificity and detail through the use of the accompanying drawings, in which:

[0013] Figure 1 A computer architecture is shown that facilitates storing a snapshot of memory reorder hints and / or processor state into a processor trace;

[0014] Figure 2A A first example of the life cycle of a machine code block and the corresponding memory value is shown;

[0015] Figure 2B A second example of a machine code block and corresponding memory value lifecycle that may present a replay challenge is shown;

[0016] Figure 2C A third example of a machine code block and corresponding memory value lifecycle that may present a replay challenge is shown; and

[0017] Figure 3A flow chart of an example method for storing memory reorder hints in a processor trace is shown. DETAILED DESCRIPTION

[0018] The inventors have recognized that there are several ways to reduce the amount of processing required to determine which of multiple recorded cache values ​​presents the correct execution result during trace replay. The first approach is for the processor designer to change the processor design so that the processor engages in less memory reordering and / or speculative execution behavior - thereby reducing the number of out-of-order cache inflows that need to be considered. The inventors have recognized that this solution may be impractical for many modern processors because memory reordering and / or speculative execution behavior have a significant impact on the performance of modern processors. As a compromise, the processor designer may cause the processor to engage in less memory reordering and / or speculative execution behavior only when the processor's trace function is enabled. However, doing so may change the way code is executed when tracing is enabled, relative to the way the same code is executed when tracing is disabled. In turn, this may change the way programming errors are indicated when tracing is enabled versus disabled.

[0019] A second approach is for the processor designer to provide additional documentation about the memory reordering and / or speculative execution behavior of the processor. With the availability of more detailed documentation, the author of time travel tracing software may be able to identify, for a given memory access instruction, a reduced number of logged cache values ​​that may render the correct execution result. As a result, this additional documentation can be used to reduce the search space for identifying the correct logged cache values. However, processor designers may be reluctant to provide additional documentation about the memory reordering and / or speculative execution behavior of the processor. For example, processor designers may want to avoid guaranteeing specific behaviors so that they have the flexibility to change these behaviors in future processors. Processor designers may even want to treat these behaviors as trade secrets.

[0020] The third method involves retaining the reordering and / or speculative execution behavior of the processor, but it modifies the processor so that when the processor performs observable memory reordering behavior, the processor provides additional trace data. Such additional trace data may be, for example, a hint about the actual operation of the processor. More particularly, in this third method, when the processor performs observable memory reordering behavior, the processor can record additional information into the processor, and the additional information can be used later to identify the specific cache value actually consumed by a given instruction. According to this third method, at least some embodiments described herein include processor modifications that cause the processor that records the execution of the thread into the trace to also record additional memory reordering hints into the trace. These hints provide information available during trace replay to help identify which memory value is actually used by a given memory access machine code instruction. Examples of memory reordering hints include information that helps identify how long ago a specific machine code instruction read from a specific cache line (for example, in terms of the number of processor cycles, the number of instructions, etc.), whether a specific machine code instruction reads a current value from a specific cache line, an indication of which value a specific machine code instruction reads from a specific cache, and the like.

[0021] The fourth method also involves retaining the reordering and / or speculative execution behavior of the processor, but it modifies the processor so that the processor records additional processor state, which can be used to provide additional boundaries for the mathematical problem of determining which cache value of multiple log records will present the correct execution result. According to this fourth method (which can be used alone or in combination with the third method), at least some embodiments described herein include modifications to the processor, which cause the processor to record additional processor state into the trace. Such additional processor state can include, for example, the value of at least one register, the hash of at least one register, an instruction count, at least part of the processor branch trace, and the like. Such processor state can provide additional boundaries for the mathematical problem of determining which cache value of multiple log records presents the correct execution result.

[0022] In order to implement the third and / or fourth methods described above, Figure 1 An example computer architecture 100 is shown that facilitates storing memory reorder hints and / or snapshots of processor state in a processor trace. Figure 1 Computer system 101 is shown to include (among other things) one or more processors 102 , a system memory 103 (e.g., random access memory), and a persistent storage device 104 (e.g., magnetic storage media, solid-state storage media, etc.) that are communicatively coupled using a communication bus 110 .

[0023] As shown, persistent storage 104 may store (among other things) a tracer 104a, one or more traces 104b, and an application 104c. During operation of computer system 101, processor 102 may load tracer 104a and application 104c into system memory 103 (i.e., tracer 103a and application 103b shown). In an embodiment, processor(s) 102 execute machine code instructions of application 103b, and during execution of those machine code instructions, tracer 103a causes processor(s) 102 to record a bit-accurate trace of the execution of those instructions. The bit-accurate trace may be recorded based at least in part on recording cache flows to cache(s) 107 (discussed later) caused by the execution of those instructions. The trace may be stored in system memory 103 (i.e., trace(s) 103b shown), and, if desired, may also be persistently stored in persistent storage (e.g., as shown by the arrow between trace(s) 103b and trace(s) 104b).

[0024] Figure 1 Certain components of each processor 102 that may implement embodiments herein are detailed. As shown, each processor 102 may include (among other components) one or more processing units 105 (e.g., processor cores), one or more caches 107 (e.g., a level 1 cache, a level 2 cache, etc.), multiple registers 108, and microcode 109. Typically, each processor 102 loads system code instructions (e.g., of an application 103c) from system memory 103 into cache(s) 107 (e.g., into the "code" portion of cache(s) 107), and executes those machine code instructions using one or more of the processing units 105. During execution of the machine code instructions, the machine code instructions may use registers 108 as temporary storage locations, and may read and write to various locations in system memory 103 via cache(s) 107 (e.g., using the "data" portion of cache(s) 107). Although the operation of the various components of each processor 102 is largely controlled by physical hardware-based logic (eg, implemented using transistors), the operation of the various components of each processor 102 may also be controlled, at least in part, using software instructions contained in processor microcode 109 .

[0025] As shown, each processing unit 105 includes a plurality of execution units 106. These execution units 106 may include, for example, arithmetic logic units, memory units, etc. In modern processors, each processing unit 105 may include multiple types of execution units, and these execution units 106 are arranged in a manner to support parallel execution of machine code instructions. In this way, each processing unit 105 can work to execute multiple machine code instructions (e.g., from application 103c) in parallel. Each processing unit 105 can be regarded as comprising a processing "pipeline" that can continuously (or periodically) receive an influx of new machine code instructions. The size and length of the pipeline (e.g., how many instructions can be processed at a time, and how long it takes to complete the execution of each instruction) is at least partially defined by the number, identification, and arrangement of the execution units 106.

[0026] In some processor implementations, parallel execution of machine code instructions is accomplished by processing unit 105 loading a sequence of machine code instructions (e.g., a fixed number of bytes) from cache(s) 107 and decoding these machine code instructions into micro-operations (μops), which are executed on execution unit 106. After decoding the instructions into μops, processing unit 105 dispatches these μops for execution on execution unit 106. Other processor implementations may execute machine code instructions directly on execution unit 106 without decoding them into μops.

[0027] Figure 2A An example 200a is shown that includes a series of machine code instructions that may be loaded from cache(s) 107 and decoded into μops by processing unit 105. These machine code instructions include a plurality of load instructions (e.g., instructions 1-4, 6, and 7)—each of which loads a value from a memory location (memory location AD) into a register (registers R1-R6)—and a store instruction (e.g., instruction 5) that stores a value contained in a register (e.g., register R1) into a memory location (e.g., memory location D). Processing unit 105 may decode each of these loads and store them into executable μops at execution unit 106.

[0028] During the decoding and / or dispatching of a given series of machine code instructions, processing unit 105 may identify multiple machine code instructions in the series that lack dependencies on each other (e.g., independent loads or stores, independent math operations, etc.), and dispatch their μops for parallel execution at execution unit 106. Additionally, or alternatively, processing unit 105 may identify machine code instructions that may or may not be executed based on the outcome of a branch / condition, and choose to speculatively schedule μops for parallel execution at execution unit 106.

[0029] For example, refer again to Figure 2A , processing unit 105 may determine that the loads of instructions 1-4 are not dependent on each other (i.e., they do not have to be executed serially as if they appeared in a series of instructions). As such, processing unit 105 may dispatch μops corresponding to instructions 1-4 for execution in parallel at execution unit 106. These μops may then proceed to execute the loads. This may include, for example, execution unit 106 identifying one or more memory values ​​already stored in cache(s) 107, execution unit 106 initiating one or more cache misses if the storage location requested by the load is not already stored in cache 107, and so on. Notably, the amount of time (e.g., processor clock cycles) it takes for a μop for a given instruction to execute may vary depending on the state of processor 102 (e.g., existing μops executing at execution unit 106, existing contents of cache(s) 107, concurrent activity of other processing units 105, and so on). For example, even though the μops corresponding to instructions 1-4 may initially be dispatched at the same time, the μops for each instruction may not complete at the same time.

[0030] Eventually, the execution unit 106 may complete the μops for a given instruction, and the processing unit 105 may "commit" (sometimes referred to as "retiring") the instruction. Most processing units commit instructions in the order in which they were originally specified in the code of the application 104c, regardless of the order in which the μops for those instructions were dispatched and / or completed. Thus, from an architectural perspective, it may appear that the processing unit 105 has executed a series of instructions in order, even though those instructions may have been executed out of order within the processing unit 105. An instruction is often referred to as being "in flight" from the time its execution is initiated at the execution unit 106 to the time it is committed.

[0031] If processing unit 105 engages in speculative execution, it may initiate the execution of instructions that may not actually need to be executed. Figure 2AThe instructions may be instructions that should only be executed if the conditions of a previous branch instruction are met. In order to keep the execution unit 106 "busy", the processing unit 105 may still initiate their execution at the execution unit 106 before the branch instructions commit. If it is later discovered that the conditions have been met (e.g., when the branch instruction commits), these speculatively executed instructions can commit when their μops are completed. On the other hand, if it is later discovered that the conditions have not been met, the processing unit 105 may avoid committing these speculatively executed instructions. In the second case, the execution of the μops for these speculatively executed instructions may have an observable effect on the processor state even though these instructions are not committed. For example, these μops may also cause cache misses - resulting in data being brought into (or out of) the (multiple) caches 107 - even if these cache misses are not ultimately consumed (e.g., because the speculative instructions that caused the cache misses were not committed).

[0032] As a side effect of out-of-order and / or speculative execution, the order in which data flows into cache(s) 107 (e.g., cache misses) may lack correspondence with the order in which execution of machine code instructions occurs. Additionally, once data is in cache(s) 107, the lifecycle of that data may lack correspondence with the cache lifecycle that might be expected based on the order in which the machine code instructions appear to have executed. If unused cache misses from speculative execution are logged to trace(s) 103b, these cache misses will add log data that is ultimately not required for the correct reply of trace(s) 103b, but this may lead to ambiguity about which value a given instruction actually reads. Further complications may arise due to concurrent execution of other threads, as these threads may further cause cache misses, exits, and invalidations. As previously described, all of this means that when replaying a trace based on recorded cache misses, additional processing may be required to determine which logged cache(s) values ​​a given machine code instruction actually read.

[0033] In addition to cache inflows, some cache-based trace logging implementations may also log cache exits and / or invalidations. Thus, trace(s) 104b may contain information sufficient to determine when a cache line was initially introduced into cache(s) 107, and when the cache line was later exited from cache(s) 107 or invalidated within cache(s) 107. This means that trace replay software can determine the total lifecycle of a particular cache line within cache(s) 107.

[0034] For example, Figure 2AAn example 200a of a possible life cycle of a cache line corresponding to memory locations A, B, and C is shown. Figure 2A As shown by the arrows in , in example 200a, these cache lines are all brought into cache(s) 107 before the load at instruction 1 is committed (e.g., because these loads of instructions 1-4 are all initiated in parallel). Figure 2A As shown by the arrows in , in example 200a, the cache line for memory location A remains valid in the cache until the load at instruction 7 is committed, the cache line for memory location B remains valid in the cache until the load at instruction 3 is committed, and the cache line for memory location C remains valid in the cache until the load at instruction 5 is committed. Through the life cycle of these examples, the trace replay software can easily identify which logged cache values ​​correspond to related loads (e.g., the loads at instructions 1-4, which read memory locations AC). This is because, when each related load instruction is committed, there is a current valid cache line that the load has read from.

[0035] Despite Figure 2A The reordering in Example 200a may not present a significant replay challenge, but Figure 2B and 2C Examples 200b, 200c are shown where reordering may present replay challenges. Figure 2B Example 200b includes Figure 2A Example 200a of Example 200b uses the same series of machine code instructions, and the same cache line lifecycle for cache lines corresponding to memory locations A and B. However, unlike Example 200a, in Example 200b, the cache line for memory location C only remains valid in the cache until the load at instruction 2 is committed. This means that the cache line corresponding to the storage location being read from the load is invalidated or retired before the load at instruction 4 commits. When the replay software considers the load at instruction 4 during replay, its invalidation / retirement before the load commits (e.g., in-flight) may present challenges. For example, the replay software may need to determine whether it is legal to read from this cache line for the load at instruction 4, even though the cache line was invalidated / retired before the load commits.

[0036] Figure 2C Example 200c also includes Figure 2AExample 200a shows the same series of machine code instructions, and the same cache line lifecycle for cache lines corresponding to storage locations A and B. However, in Example 200c, the cache line at memory location C is valid in cache 107 before the load at instruction 1 to the load commit at instruction 2, and then from the commit of the load at instruction 3 to the commit of the load at instruction 6. Assuming that the value of memory location C is the same during these two validity periods, the replay software does not need to distinguish between the two validity periods when considering the load at instruction 4. However, if the values ​​within the two validity periods are different, the replay software will need to determine which of the two values ​​was actually read by the load at instruction 4. Therefore, when the replay software considers the load at instruction 4 during replay, the change in the value of the cache line before the load commits (e.g., while it is processing in-flight) can also pose challenges.

[0037] To assist the replay software in determining which cache lines are valid for a load and / or which cache value was read, embodiments include processor modifications that detect situations where out-of-order and / or speculative execution may result in observable effects and therefore store one or more memory reorder hints as a result. These processor modifications are described in Figure 1 109, but it should be understood that these processor modifications may potentially be implemented as physical logic changes in addition to (or in lieu of) microcode 109 changes.

[0038] In general, the reorder hint logic 109a detects the following situations: (i) the execution of a machine code instruction that performs a load from a memory address is initiated (e.g., its μops are distributed to the execution unit 106); (ii) the execution of the machine code instruction causes a particular cache line in the cache(s) 107 (e.g., a cache line that overlaps the memory address) to be logged to the trace(s) 103; (iii) after the particular cache line is logged (but before the machine code instruction is committed), an event is determined to have affected a particular cache line in the cache. For example, this may occur due to the following reasons: a cache line is retired or invalidated, a cache line is written, a read lock on a particular cache line is lost, etc. When this occurs, the reorder hint logic 109a may cause the processor 102 to log one or more memory reorder hints to the trace(s) 103b. In general, the memory reorder hints may include any data that can be used to help the replay software identify which cache lines are valid for the machine code instruction and / or which value was read by the machine code instruction. Notably, if an instruction never commits (eg, because it was speculatively executed and later discovered not to be needed), the reorder hint logic 109a may refrain from recording any hints for that instruction.

[0039] In an embodiment, the reorder hint logic 109a may record memory reorder hints only when memory access behavior has deviated from a defined "normal" behavior. For example, the processor 102 may define the normal behavior as: instructions generally use the value in the cache(s) 107 when the instruction commits. The reorder hint logic 109a may then record a reorder hint only when an instruction uses (or may have used) a value that is different from the value in the cache(s) 107 when the instruction commits. In other words, the reorder hint logic 109a only records a reorder hint when an instruction uses an "old" value (e.g., a cache line from the log record above), not when a "new" or "current" value (e.g., resulting from a subsequent cache invalidation / retirement or cache line write) is used by the instruction. In this way, a reorder hint is recorded only when there is a deviation from normal behavior and the absence of a reorder hint contains implicit knowledge (e.g., a current value is used). Thus, in these embodiments, a reorder hint may be stored only when an event affecting a particular cache line changes its value, and when the processor uses an old value.

[0040] For example, return to Figure 2B , if the reorder hint logic 109a were to detect that the cache line corresponding to memory address C was invalidated / retired before the load at instruction 4 was committed, the reorder hint logic 109a might record an indication of how long ago the load at instruction 4 was read from the cache line, or how long ago the load at instruction 4 might have been read from the cache line (e.g., how long the processor pipeline / read-ahead window is). This could, for example, be expressed in terms of a number of processor cycles, a number of instructions, etc. Then returning to Figure 2C , if the reorder hint logic 109a were to detect that the memory value corresponding to memory address C had changed prior to the commit of the load at instruction 4, the reorder hint logic 109a could record an indication of which memory value was read. This could be expressed as how long ago the load at instruction 4 read from the cache line, whether the load at instruction 4 read the current value from the cache line, or even what value the load at instruction 4 read.

[0041] In some embodiments, the reorder hint logic 109a may be probabilistic. For example, if the reorder hint logic 109a records a hint when an instruction uses a value other than the current value when the instruction is submitted, the reorder hint logic 109a may also record a hint when it is uncertain whether the instruction uses the current value when the instruction is submitted. In addition, the reorder hint logic 109a may even record the probability that the most recent value is used (or not used). Recording the probability allows the replay software to narrow the search space by testing the most likely path first. For example, a probability of 0% may mean that the reorder hint logic 109a can determine that the previous value is used, while a probability of 100% (or if the data is implicit, there is no data packet) may mean that the reorder hint logic 109a can determine that the current value is used. Values ​​between 0% and 100% can specify the probability of using the current value. Different numbers of bits can be used for different probability granularities. For example, {100%, 0%} for one bit, {100%, 99-50%, 49-1%, 0%} for two bits, {100%, 99-90%, 89-51%, 50%, 49-10%, 9-1%, 0%} for three bits, {100% = 0xFF, 0% = 0x00, other values, probability*256}8 bits, floating point value, etc.

[0042] In view of the foregoing, Figure 3 A flow chart of an example method 300 for storing memory reorder hints in a processor trace is shown. In general, the method 300 is implemented at a computing system (e.g., processor 102) that includes one or more processing units (e.g., processing unit 105) and a processor cache (e.g., cache(s) 107) that includes a plurality of cache lines. The method 300 may be implemented in an environment where the processor does not typically log processor state (e.g., registers such as instruction pointers, instruction counts, etc.). Reference will be made to Figure 1 Computer Architecture 100 and Figure 2B and 2C Method 300 is described with reference to examples 200b and 200c.

[0043] like Figure 3As shown, method 300 includes an action 301 of initiating execution of a load. In some embodiments, action 301 includes initiating execution of a specific machine code instruction while executing multiple machine code instructions at one or more processing units, the specific machine code instruction performing a load to a memory address. For example, while tracking execution of a thread of application 104c at processing unit 105, processing unit 105 may fetch a series of instructions of application 104c from cache(s) 107, such as the series of instructions shown in examples 200b and 200c. Processing unit 105 may then decode one or more of these instructions into multiple μops and dispatch these μops for execution at execution unit 106. As part of this fetch / decode process, processing unit 105 may decode the load at instruction 4 and dispatch μops for execution at execution unit 106.

[0044] Method 300 also includes an act 302 of logging a cache line as a result of the load. In some embodiments, act 302 includes, based on initiation of a particular machine code instruction, initiating logging of a particular cache line in a processor cache that overlaps a memory address in a processor trace, including initiating logging of a value corresponding to the memory address associated with logging the particular cache line. For example, based on execution of the μops for load at instruction 4 at execution unit 106, execution unit 106 may cause a cache miss based on accessing a memory address used by the load (e.g., address C), resulting in a value in system memory 103 corresponding to memory address C being loaded into a cache line in cache(s) 107. Alternatively, the μops for load of instruction 4 may read a value from a cache line already present in cache(s) 107, which value is then logged into trace(s) 103b as a result of the load. The logging may occur in association with the load, or at other times. Figure 2B The life cycle of this cache line can be represented by the arrow corresponding to the memory address C, and Figure 2C , the life cycle of this cache line can be represented by one of the arrows corresponding to the memory address C.

[0045] Method 300 also includes an act 303 of detecting an event affecting a cache line before committing the load. In some embodiments, act 303 includes detecting an event affecting a particular cache line after initiating logging of the particular cache line to a processor trace and before committing the particular machine code instruction. For example, Figure 2BIn the context of , the reorder hint logic 109a (e.g., microcode and / or physical logic) can detect that the cache line corresponding to memory address C is retired or invalidated before the load at instruction 4 is committed (and, therefore, there is an invalidation or retirement of a particular cache line before a particular machine code instruction is committed). In another example, Figure 2C In the context of , the reorder hint logic 109a (e.g., microcode and / or physical logic) can detect that: (i) the first cache line corresponding to memory address C is retired / invalidated before the load at instruction 4, and (ii) a new cache line with a different value for memory address C is brought into the cache before the load at instruction 4 commits (and, therefore, there is a change in the particular cache line, which includes a change in the storage value corresponding to the memory address before the particular machine code instruction is committed). Other events may include writes to the particular cache line, loss of a read lock on the particular cache line, etc., which may cause the value of the cache line to change after it is recorded. These situations may be caused by speculative execution, activity of other threads, etc.

[0046] Method 300 also includes an act 304 based on detecting a storage reorder hint. In some embodiments, act 304 includes initiating storage of a memory reorder hint to a processor trace based at least on detecting an event affecting a particular cache line. For example, in Figure 2B In the context of , the reorder hint logic 109a may store a hint in the trace(s) 103b of how long ago a particular machine code instruction was read from a particular cache line or how long ago a particular machine code instruction has been read from a particular cache line. As discussed, this may be expressed in terms of a number of processor cycles, a number of instructions, etc., which describes the size of the processor's read-ahead window. Figure 2C In the context of , the reorder hint logic 109a may store in the trace(s) 103b an indication of which memory value was read, for example, whether a particular machine code instruction read the current value from a particular cache line, or which value a particular machine code instruction read from a particular cache line.

[0047] As mentioned, some embodiments may only log memory reorder hints if memory access behavior has deviated from a defined general behavior. For example, if the general behavior is that instructions typically use the value in the cache(s) when the instruction commits, then act 304 may initiate the storage of the memory reorder hint into the processor trace only when the machine code instruction loads the value logged in act 302 (e.g., when it does not load a new value resulting from an event affecting a particular cache line, so it loads the old value).

[0048] As previously described, bit-accurate traces can record not only cache inflows, but also cache exits and / or invalidations. Thus, method 300 can include initiating storage of a record of at least one of the following: a later invalidation of a particular cache line, or a later exit of a particular cache line (e.g., recording the corresponding cache line exits / invalidations in trace(s) 103b) into a processor trace. These records can be used during replay to identify the value brought into cache(s) 107 and its lifecycle. The lifecycle information can be combined with reorder hints to help identify which particular cache value a particular instruction reads.

[0049] Thus, an embodiment may include a processor modification that causes the processor to record memory reordering hints into the trace. These hints provide information that can be used during trace replay to help identify which memory value was actually used by a given memory access machine code instruction. These hints can significantly reduce the processing required to perform trace replay.

[0050] As mentioned, some embodiments may additionally (or alternatively) record additional processor states into the trace(s) 103b. These states may provide additional information about the state of the program, adding bounds to the mathematical problem of determining which of the multiple recorded cache values ​​presents the correct execution result. Thus, embodiments may include processor modifications that periodically or continuously record additional states into the trace(s) 103b. These processor modifications may be used to record additional states into the trace(s) 103b. Figure 1 109. However, similar to the reorder hint logic 109a, it should be understood that these processor modifications may also potentially be implemented as physical logic modifications in addition to (or instead of) microcode 109 changes.

[0051] Processor state logic 109b may record snapshots of periodic processor state at regular intervals (e.g., based on the number of instructions executed since the last snapshot, the number of processor clock cycles that have passed since the last snapshot, etc.). In each snapshot, processor state logic 109b may record any state that may be used to help constrain the mathematical problem that determines which of the multiple logged cache values ​​will render the correct execution result. Examples of available processor state include the value(s) of one or more registers, a hash of the value(s) of one or more registers, an instruction count (e.g., of the next instruction to be executed, of the last instruction committed, etc.), and the like.

[0052] In some embodiments, processor state logic 109b may even record a more continuous stream of additional processor states. For example, many modern processors include functionality for generating a "branch trace", which is a trace that indicates which branches were taken / not taken when executing code. Examples of branch trace technology include INTELPROCESSOR TRACE and ARM PROGRAM TRACE MACROCELL. When branch trace is available, processor state logic 109b may record all or a subset of this branch trace into (multiple) traces 103b. For example, processor state logic 109b may record the entire branch trace (e.g., as a separate data stream in (multiple) traces 103b), processor state logic 109b may record a sampling of branch traces (e.g., the result of each indirect jump, plus the number of jumps defined after the indirect jump, or the number of bytes of branch trace data defined after the indirect jump), and / or processor state logic 109b may record a subset of branch traces (e.g., only the results of indirect jumps).

[0053] In view of the foregoing, it will be appreciated that method 300 may also include initiating storage of additional processor state into a processor trace. The processor state may include, for example, one or more of the following: a value of at least one register, a hash of at least one register, an instruction count, and / or at least a portion of a branch trace. This data may be recorded as periodic snapshots or as a more continuous data stream. If the processor state includes a portion of a branch trace, the processor state may include a sample of at least one branch trace or a subset of the branch trace.

[0054] Thus, embodiments may also include a modification of the processor that causes the processor to record additional processor state into the trace. Such additional processor state may include a snapshot of register values, a hash of register values, an instruction count, etc. Such additional processor state may additionally or alternatively include at least a portion of a processor branch trace. This recorded processor state may provide additional boundaries for the mathematical problem of determining which of a plurality of logged cache values ​​will present the correct execution result, thereby reducing the processing required to perform a trace replay.

[0055] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or acts described above, or the order of the acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0056] Embodiments of the present invention may include or utilize a special or general-purpose computer system including computer hardware, such as, for example, one or more processors and system memory, as discussed in more detail below. Embodiments within the scope of the present invention also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media may be any available media that can be accessed by a general or special-purpose computer system. The computer-readable medium that stores computer-executable instructions and / or data structures may be a computer storage medium. The computer-readable medium that executes computer-executable instructions and / or data structures is a transmission medium. Therefore, by way of example and not limitation, embodiments of the present invention may include at least two distinct types of computer-readable media: a computer storage medium and a transmission medium.

[0057] Computer storage media are physical storage media that store computer-executable instructions and / or data structures. Physical storage media include computer hardware, such as RAM, ROM, EEPROM, solid-state drives ("SSD"), flash memory, phase-change memory ("PCM"), optical disk storage, magnetic disk storage or other magnetic storage devices, or any other hardware storage device(s) that can be used to store program code in the form of computer-executable instructions or data structures that can be accessed and executed by a general-purpose or special-purpose computer system to implement the functions disclosed in the present invention.

[0058] Transmission media may include networks and / or data links that can be used to carry program code in the form of computer executable instructions or data structures and can be accessed by general or special computer systems. "Network" is defined as one or more data links that enable electronic data to be transmitted between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer system via a network or another communication connection (hardwired, wireless, or a combination of hardwired or wireless), the computer system may view the connection as a transmission medium. The above combination should also be included in the scope of computer-readable media.

[0059] In addition, program code in the form of computer executable instructions or data structures may be automatically transferred from transmission media to computer storage media (or vice versa) upon reaching various computer system components. For example, computer executable instructions or data structures received over a network or data link may be buffered in RAM within a network interface module (e.g., "NIC") and then ultimately transferred to computer system RAM and / or to less volatile computer storage media at the computer system. Thus, it should be understood that computer storage media may be included in computer system components that also (or even primarily) utilize transmission media.

[0060] Computer executable instructions include, for example, instructions and data that, when executed on one or more processors, cause a general purpose computer system, a special purpose computer system, or a special purpose processing device to perform a certain function or group of functions. Computer executable instructions may be, for example, binary, intermediate format instructions such as assembly language, or even source code.

[0061] Those skilled in the art will appreciate that the present invention can be realized in a network computing environment with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, based on microprocessor or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, etc. The present invention can also be realized in a distributed system environment, in which local and remote computer systems that are linked by a network (by a hardwired data link, a wireless data link, or by a combination of a hardwired data link) all perform tasks. Like this, in a distributed system environment, a computer system can include multiple component computer systems. In a distributed system environment, a program module can be located in a local and remote storage device.

[0062] Those skilled in the art will also appreciate that the present invention may be implemented in a cloud computing environment. A cloud computing environment may be distributed, although this is not required. When distributed, a cloud computing environment may be internationally distributed among organizations and / or have components owned across multiple organizations. In this specification and the appended claims, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services). The definition of "cloud computing" is not limited to any of the other numerous advantages that may be derived from such a model when properly deployed.

[0063] Cloud computing models can include various features, such as on-demand self-service, broad network access, resource pooling, rapid elasticity, measured services, etc. Cloud computing models can also take the form of various service models, such as software as a service ("SaaS"), platform as a service ("PaaS"), and infrastructure as a service ("IaaS"). Cloud computing models can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, etc.

[0064] Some embodiments, such as cloud computing environments, may include a system that includes one or more hosts, each of which is capable of running one or more virtual machines. During operation, the virtual machine will simulate a running computing system, thereby supporting an operating system and possibly one or more other applications. In some embodiments, each host includes a hypervisor that simulates the virtual resources of the virtual machine using physical resources abstracted from the perspective of the virtual machine. The hypervisor can also provide appropriate isolation between virtual machines. Thus, from the perspective of any given virtual machine, the hypervisor provides the illusion that the virtual machine is interfacing with physical resources, even though the virtual machine is only interacting with the appearance of physical resources (e.g., virtual resources). Examples of physical resources include processing power, memory, disk space, network bandwidth, media drives, etc.

[0065] The present invention may be embodied in other specific forms without departing from the spirit or essential features of the present invention. The described embodiments are considered in all respects to be illustrative only and not restrictive. Therefore, the scope of the present invention is indicated by the appended claims rather than the preceding description. All changes falling within the equivalent meaning and scope of the claims should be included within their scope.

Claims

1. A system for storing memory reorder hints in a processor trace record, the system comprising: one or more processing units; a processor cache, comprising a plurality of cache lines; as well as A logic module configured to perform at least the following operations while executing a plurality of machine code instructions at the one or more processing units: initiating execution of a machine code instruction that performs a load of a memory address; initiating, based on the initiation of the machine code instruction, logging a cache line in the processor cache that overlaps the memory address into the processor trace record, including initiating logging of a value corresponding to the memory address in connection with logging the cache line; detecting an event affecting the cache line after initiating logging of the cache line to the processor trace record and before committing the machine code instructions; as well as Based at least on detecting the event affecting the cache line, storage of a memory reorder hint into the processor trace record is initiated.

2. The system of claim 1 , wherein the event affecting the cache line is selected from the group consisting of: an invalidation of the cache line, a retirement of the cache line, a write to the cache line, or a loss of a read lock on the cache line.

3. The system of claim 2, wherein the event comprises an invalidation or retirement of the cache line.

4. The system of claim 2, wherein the event comprises the write to the cache line.

5. The system of claim 1, wherein the system initiates storage of the memory reorder hint into the processor trace record only when the machine code instruction loads the value logged in relation to the cache line.

6. The system of claim 1 , wherein the memory reorder hint comprises at least one of: how long ago the machine code instruction was read from the cache line; whether the machine code instruction reads a current value from the cache line; or The machine code instruction provides an indication of which value to read from the cache.

7. The system of claim 6, wherein the memory reorder hint comprises how long ago the machine code instruction was read from the cache line, and wherein the memory reorder hint comprises at least one of: a number of processor cycles or a number of instructions.

8. The system of claim 1, wherein the system initiates storage of a read-ahead window size into the processor trace record.

9. The system of claim 1 , wherein the system further initiates storing a record of at least one of the following into the processor trace record: A later invalidation of the cache line; or The cache line is later retired.

10. The system of claim 1, wherein the system further initiates storage of additional processor state into the processor trace record.

11. The system of claim 10, wherein the processor state comprises at least one of: the value of at least one register; a hash of at least one register; instruction count; or At least part of the result of branch tracing.

12. The system of claim 11, wherein the processor state comprises the portion of the result of the branch trace, and wherein the portion of the result of the branch trace comprises at least one of: a sample of the branch trace or a subset of the branch trace.

13. The system of claim 11, wherein the processor state comprises the portion of the result of the branch trace, and wherein the branch trace comprises at least one of: Intel processor trace or ARM program trace macrocell.

14. The system of claim 1, wherein the logic module comprises processor microcode.

15. A method implemented at a computing system, the computing system comprising one or more processing units and a processor cache, the processor cache comprising a plurality of cache lines, the method for storing a memory reorder hint in a processor trace record, the method comprising: while executing a plurality of machine code instructions at the one or more processing units, initiating execution of a machine code instruction that performs a load to a memory address; initiating, based on the initiation of the machine code instruction, logging a cache line in the processor cache that overlaps the memory address into the processor trace record, including initiating logging of a value corresponding to the memory address in connection with logging the cache line; detecting an event affecting the cache line after initiating logging of the cache line to the processor trace record and before committing the machine code instructions; as well as Based at least on detecting the event affecting the cache line, storage of a memory reorder hint into the processor trace record is initiated.

16. The method of claim 15, wherein the event affecting the cache line is selected from the group consisting of: invalidation of the cache line, retirement of the cache line, a write to the cache line, or loss of a read lock on the cache line.

17. The method of claim 15, wherein the memory reorder hint comprises at least one of: how long ago the machine code instruction was read from the cache line; whether the machine code instruction reads a current value from the cache line; or The machine code instruction provides an indication of which value to read from the cache.

18. The method of claim 17, wherein the memory reorder hint comprises how long ago the machine code instruction was read from the cache line, and wherein the memory reorder hint comprises at least one of: a number of processor cycles or a number of instructions.

19. The method of claim 15, wherein the system periodically initiates storage of a processor state into the processor trace record, the processor state comprising at least one of: a value of at least one register; a hash of at least one register; an instruction count; or at least a portion of a result of a branch trace.

20. A microprocessor comprising: one or more processor cores; a cache memory including a plurality of cache lines; as well as A processor microcode that performs at least the following operations: executing a plurality of machine code instructions at the one or more processor cores; as well as while executing the plurality of machine code instructions, initiating execution of a machine code instruction that performs a load of a memory address; initiating, based on initiation of the machine code instruction, logging a cache line in the cache that overlaps the memory address into the processor trace record, including initiating logging of a value corresponding to the memory address associated with logging the cache line; detecting an event affecting the cache line after initiating logging of the cache line to a processor trace record and before committing the machine code instructions; as well as Based at least on detecting the event affecting the cache line, storage of a memory reorder hint into the processor trace record is initiated.