Computer system, method, and hardware storage device for recording a replayable trace

By adopting the identification tracking memory model and executing multi-threading methods in multi-threading programs, the problems of low performance and excessive tracking files in the prior art are solved, and efficient debugging recording and playback in multi-core and hyperthreading environments are achieved.

CN114490292BActive Publication Date: 2025-07-29MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210093599.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-08-31
Filing Date
2017-08-23
Publication Date
2025-07-29
Estimated Expiration
2037-08-23

AI Technical Summary

Technical Problem

The existing time travel debugging technology has problems such as poor performance and excessive tracking of files in multi-threaded programs, and cannot effectively support debugging requirements in multi-core and hyperthreaded environments.

Method used

Using the identification tracking memory model, multiple threads sort orderable events across multiple threads of multi-threaded processes, and multiple threads are executed simultaneously on multiple processors, independently recording the playbackable trace of each thread, including recording the initial state of the thread, the side effects of non-deterministic instructions and memory reading, and using monotonously increased numerical sorting events to reduce the size of the tracking file.

Benefits of technology

It realizes high-performance recording and playback in multi-threaded programs, reduces the size of tracking files, supports debugging requirements in multi-core and hyper-threading environments, and improves debugging efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114490292B_ABST
    Figure CN114490292B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to computer systems, methods, and hardware storage devices for recording replayable traces. Recording a replayable trace of the execution of a multithreaded process includes identifying a trace memory model that defines one or more sortable events that are to be ordered across multiple threads of the multithreaded process. Multiple threads are executed simultaneously across one or more processing units of one or more processors. During the execution of the multiple threads, a separate replayable trace is recorded independently for each thread. The recording includes, for each thread: recording an initial state for the thread; recording at least one memory read performed by at least one processor instruction, where the at least one processor instruction executed by the thread takes the memory as an input; and recording at least one sortable event executed by the thread with a monotonically increasing number that orders the event among other sortable events across the multiple threads.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with the international filing date of August 23, 2017, international application number PCT / US2017 / 048094, which entered the Chinese national phase on February 27, 2019, Chinese national application number 201780053059.9, and the invention title of "Computer System, Method, and Hardware Storage Device for Recording Replayable Traces". Technical Field

[0002] Embodiments of the present disclosure relate to the field of computers, and more particularly to recording replayable traces of the execution of multi-threaded processes. Background Art

[0003] When writing source code during the development of a software application, developers typically spend a significant amount of time "debugging" the source code to find runtime errors in the code. For example, developers can take several approaches to reproduce and locate source code defects, such as observing the behavior of the program based on different inputs, inserting debugging code (e.g., to print variable values, trace executed branches, etc.), temporarily removing code sections, etc. Tracking runtime errors to identify code defects accounts for a significant portion of the application development time.

[0004] Debugging applications ("debuggers") have been developed to assist in the code debugging process. Many such tools provide the ability to trace, visualize, and change the execution of computer code. For example, a debugger can visualize the execution of code instructions (e.g., source code, assembly code, etc.) and variable values, and enable the user to change aspects of the code execution. Typically, a debugger enables the user to set "breakpoints" in the source code (e.g., specific instructions or statements in the source code), which cause the execution of the program to be suspended when reached during execution. When the source code execution is suspended, the user can be presented with variable values and given options on how to proceed (e.g., by terminating the execution, by continuing execution as normal, by stepping into a statement / function call, stepping through a statement / function call, or stepping out of a statement / function call, etc.). However, classical debuggers only enable the code execution to be observed / changed in a single direction (forward). For example, classical debuggers do not enable the user to select to return to a previous breakpoint.

[0005] An emerging form of debugging is "time travel" debugging, in which the execution of a program is recorded into a trace that can later be replayed and analyzed both forward and backward. Summary of the Invention

[0006] Embodiments herein relate to a new implementation of recording and replaying traces for time-travel debugging, which can produce performance improvements by several orders of magnitude over previous attempts, which implements recording of multi-threaded programs where threads are free to run simultaneously across multiple processing units, and which can produce trace files having a size reduced by several orders of magnitude compared to previously attempted trace files.

[0007] In some embodiments, a method for recording a replayable trace of the execution of a multi-threaded process includes identifying a trace memory model that defines one or more sortable events that are to be ordered across multiple threads of the multi-threaded process. The method further includes simultaneously executing multiple threads across one or more processing units of one or more processors, and during the execution of the multiple threads, independently recording a separate replayable trace for each thread. Recording a separate replayable trace for each thread includes recording an initial state for the thread. Recording a separate replayable trace for each thread further includes: recording at least one memory read performed by at least one processor instruction, where the memory is taken as input by at least one processor instruction executed by the thread. Recording a separate replayable trace for each thread further includes recording at least one sortable event executed by the thread with a monotonically increasing number that orders the event among other sortable events across multiple threads.

[0008] This summary is provided to introduce a series of concepts that are further described below in the detailed description in a simplified form. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To describe the manner in which the above-recited and other advantages and features of the invention can be obtained, a more particular description of the invention briefly described above will be rendered by reference to specific embodiments thereof that are illustrated in the accompanying drawings. In the understanding that the drawings illustrate only typical embodiments of the invention and are not to be considered limiting of its scope, the invention will be described and explained with additional specificity and detail in the use of the drawings, in which:

[0010] Figure 1 An example computer architecture in which embodiments of time-travel debugging can operate is illustrated;

[0011] Figure 2 A flowchart of an example method for recording a replayable trace of the execution of a multi-threaded process is illustrated;

[0012] Figure 3 is an example of sorted events across simultaneously executing threads;

[0013] Figure 4 An example of the use of shadow copies is illustrated;

[0014] Figure 5 An example of a circular buffer is illustrated;

[0015] Figure 6 An example computer architecture for processor cache-based tracing is illustrated; and

[0016] Figure 7 A flowchart of an example method for using cache data to record a replayable trace of the execution of an executable entity is illustrated. DETAILED DESCRIPTION

[0017] Embodiments herein relate to a new implementation of recording and replaying traces for time travel debugging, which can produce performance improvements of several orders of magnitude over previous attempts, which enables recording of multithreaded programs where threads freely run simultaneously across multiple processing units, and which can produce trace files having a size reduced by several orders of magnitude compared to previously attempted trace files.

[0018] Generally, the goal of time travel debugging is to capture in a trace which processor instructions are executed by an executable entity (e.g., a user-mode thread, a kernel thread, a hypervisor, etc.), such that these instructions can be replayed at a later time with absolute precision according to the trace, regardless of what granularity is desired. Being able to replay every instruction that was executed as part of the application code gives the illusion of a backward replay of the application at a later time. For example, in order to hit a breakpoint in the backward direction, the trace is replayed from a time before the breakpoint, and the replay stops at the last time the breakpoint was hit (which is before where the debugger is analyzing the code flow).

[0019] Previous attempts at providing time travel debugging have suffered several drawbacks leading to limited adoption. For example, previous attempts imposed significant constraints on code execution, such as requiring a strict ordering of all instructions included in the execution (i.e., a totally sequentially consistent recording model). This was accomplished, for example, by requiring multi-threaded programs to be executed non-simultaneously on a single core or by requiring program instructions to be executed non-simultaneously in lockstep on multiple cores (e.g., execute N instructions on one processor core and then N instructions on another processor core, and so on). These are significant limitations given today's highly multi-threaded code and highly parallel multi-core and hyper-threaded processors. Additionally, previous attempts caused significant program performance degradation and generally produced extremely large trace files (especially when simulating multiple cores), at least in part because they deterministically record the execution of every instruction and create a comprehensive record of the full memory state during program execution. Each of the foregoing has made previous attempts at time travel debugging both extremely slow and impractical for use in a production environment and for long-term tracing of program execution (especially for applications with many threads running simultaneously).

[0020] Operating environment

[0021] First, Figure 1 FIG. illustrates an example computing environment 100 in which embodiments of time travel debugging in accordance with the present invention may operate. Embodiments of the present invention may include or utilize a special purpose or general purpose computer system 101 that includes computer hardware such as one or more processors 102, system memory 103, one or more data repositories 104, and / or input / output hardware 105.

[0022] Embodiments within the scope of the present invention include physical media and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media that can be accessed by computer system 101. A computer-readable medium that stores computer-executable instructions and / or data structures is a computer storage device. A computer-readable medium that carries computer-executable instructions and / or data structures is a transmission medium. Thus, by way of example and not limitation, embodiments of the present invention may include at least two distinctly different kinds of computer-readable media: computer storage devices and transmission media.

[0023] A computer storage device is a physical hardware device that stores computer-executable instructions and / or data structures. Computer storage devices include various computer hardware, such as RAM, ROM, EEPROM, solid state drives (“SSDs”), flash memory, phase change memory (“PCM”), optical disc storage, magnetic disk storage, or other magnetic storage devices, or any other (one or more) hardware devices that can be used to store program code in the form of computer-executable instructions or data structures and that can be accessed and executed by a computer system 101 to implement the disclosed functionality of the present invention. Thus, for example, a computer storage device can include the depicted system memory 103 and / or the depicted data repository 104, which can store computer-executable instructions and / or data structures.

[0024] A transmission medium can include a network and / or a data link that can be used to carry program code in the form of computer-executable instructions or data structures and that can be accessed by a computer system 101. A “network” is defined as one or more data links that support the transfer of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided to a computer system via a network or an additional communication connection (wired, wireless, or a combination of wired or wireless), the computer system can view the connection as a transmission medium. Combinations of the above should also be included within the scope of computer-readable media. For example, input / output hardware 105 can include hardware (e.g., a network interface module (e.g., “NIC”)) that connects a network and / or a data link that can be used to execute program code in the form of computer-executable instructions or data structures.

[0025] Additionally, upon reaching various computer system components, program code in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to a computer storage device (or vice versa). For example, computer-executable instructions or data structures received via a network or a data link can be buffered in RAM within a NIC (e.g., input / output hardware 105) and then ultimately transferred to system memory 103 and / or to a non-volatile computer storage device (e.g., data repository 104) at the computer system 101. Thus, it should be understood that a computer storage device can be included in computer system components that also (or even primarily) utilize a transmission medium.

[0026] Computer-executable instructions include, for example, instructions and data that, when executed at one or more processors 102, cause a computer system 101 to perform a particular function or a group of functions. Computer-executable instructions can be, for example, binary, intermediate format instructions such as assembly language, or even source code.

[0027] Those skilled in the art will recognize that the present invention may be practiced in network computing environments having many types of computer system configurations, including personal computers, desktop computers, laptop computers, messaging processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, and the like. The present invention may also be practiced in distributed system environments where both local and remote computer systems, which are linked together by a network (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links), both execute tasks. Thus, in a distributed system environment, a computer system may include multiple constituent computer systems. In a distributed system environment, program modules may be located in both local and remote memory storage devices.

[0028] As illustrated, the data repository 104 may store computer-executable instructions and / or data structures representing the time travel debugger 106 and the application code 107, which is the object of tracing / debugging performed by the time travel debugger 106. When these programs (e.g., using the (one or more) processors 102) are executing, the system memory 103 may store corresponding runtime data, such as runtime data structures, computer-executable instructions, and the like. Thus, Figure 1 The system memory 103 is illustrated as including time travel debugger runtime data 106' and application code runtime data 107'.

[0029] As depicted, the time travel debugger 106 includes one or more components or modules, such as a recording component 106a and a replay component 106b. At appropriate times, these components may also include corresponding runtime data in the system memory 103 (illustrated as recorded runtime data 106a' and replay runtime data 106b'). During execution, the recording module 106a / recorded runtime data 106a' records one or more trace files 108 documenting the execution of the application code 107 at the (one or more) processors 102. Later, the replay module 106b / replay runtime data 106b' may use the (one or more) trace files 108 in conjunction with the application code 107 to replay the execution of the application code 107 both forward and backward. Although the (one or more) trace files 108 are depicted as being stored in the data repository 104, these (one or more) trace files 108 may also be recorded at least temporarily in the system memory 103 or at some other storage device. It should be noted that the recording component 106a may reside at one computer system, and the replay component 106b may reside at another computer system. Thus, the execution of a program may be traced / recorded on one system and replayed on another system.

[0030] Figure 1 A general representation of the internal hardware components including the (one or more) processors 102. As illustrated, the (one or more) processors 102 include one or more processing units 102a (i.e., cores). Each processing unit 102a includes hardware logic for executing processor instructions defined by an application, and these instructions are selected from a predefined processor instruction set architecture. The specific instruction set architecture of the (one or more) processors 102 varies based on the processor manufacturer and processor model. Common instruction set architectures include the IA-64 and IA-32 architectures from Intel Corporation, the AMD64 architecture from ADVANCED MICRODEVICES, Inc., and various Advanced RISC Machines (“ARM”) from ARM HOLDINGS, PLC, but numerous other instruction set architectures exist and may be used by the present invention. Generally, an “instruction” is the smallest externally visible (i.e., external to the processor) code unit that can be executed by a processor.

[0031] The processing unit 102a obtains processor instructions from the cache 102b and executes the processor instructions based on the data in the cache 102b, based on the data in the register 102c, and / or in the absence of input data. Generally, the cache 102b is a small amount (i.e., small relative to the common amount of the system memory 103) of random access memory that is a processor-on copy of a portion of the system memory 103. For example, when executing the application code 107, the cache 102b contains a portion of the application code runtime data 107'. If the processing unit(s) 102a need data that is not yet stored in the cache 102b, a "cache miss" occurs and the data is retrieved from the system memory 103 (usually evicting some other data from the cache 102b). The cache 102b is generally divided into at least a code cache and a data cache. For example, when executing the application code 107, the code cache stores at least a portion of the processor instructions stored in the application code runtime data 107', and the data cache stores at least a portion of the data structures of the application code runtime data 107'. Generally, the cache 102b is divided into separate levels (e.g., level 1, level 2, and level 3), some of which (e.g., level 3) may exist separately from the processor 102. The register 102c is a hardware-based storage location defined based on the instruction set architecture of the processor(s) 102.

[0032] Although not explicitly depicted, each processor in the processing unit(s) 102 typically includes multiple processing units 102a. Thus, a computer system can include multiple different processors 102, each of which includes multiple processing cores. In these cases, the cache 102b of the processor can include multiple distinct cache portions, each corresponding to a different processing unit, and the registers can include distinct sets of registers, each corresponding to a different processor unit. The computer system 101 can thus execute multiple "threads" simultaneously at different processors 102 and / or at different processing units 102a within each processor.

[0033] Time Travel Debugging

[0034] As previously mentioned, prior debugging in time travel debugging would execute multiple threads of a process non-simultaneously on a single processor core, or execute multiple threads non-simultaneously on different processors and / or processor cores, such that each instruction is executed and recorded in a precise, deterministic order. Additionally, prior attempts would exhaustively record changes to the memory state of a process in a deterministic manner such that each memory value is known at any given time. However, embodiments herein are capable of simultaneously executing and tracking multiple threads, thereby removing the requirement to execute and record each instruction in a precise order, and are capable of replaying while recording much less than a full record of instruction execution and memory state.

[0035] At a conceptual level, embodiments herein record a trace of the execution of one or more threads of a process individually on one or more processors, and record these one or more traces in a (one or more) trace file 108, which can be used to reproduce the inputs and outputs of each processor instruction that was executed as part of each thread (without having to record every instruction that was executed), which includes an approximation of the order of instruction execution across different threads, and which stores enough information to predict relevant memory values without having to exhaustively record the full changes to the memory state. It should be noted that although embodiments herein can trace all threads of a process, they can also trace only a subset of the threads of a process. Additionally, it should be noted that embodiments herein can trace the execution of application code 107 on (one or more) physical processors, on (one or more) virtual processors (such as processors simulated by software), and / or even in a virtual machine environment (e.g., .NET from Microsoft Corporation, JAVA from Oracle Corporation, etc.). As an example of tracing within a virtual machine environment, recording a JAVA program can include recording what operations and memory reads the "virtual processors" of the JAVA virtual machine performed. Alternatively, recording a JAVA program can include both the JAVA program and the JAVA virtual machine (e.g., by recording the execution of native code that "just-in-time" compiles JAVA code, executes a garbage collector, etc.). In the latter case, the time travel debugger 106 can be configured to separate the replay of different layers (i.e., application code versus JAVA virtual machine code).

[0036] The embodiments herein are based on the inventors' recognition that processor instructions (including virtual machine "virtual processor" instructions) can generally fall into one of three categories: (1) instructions that are labeled "non-deterministic" because they produce unpredictable outputs, as their outputs are not completely determined by the data in general registers or memory, (2) deterministic instructions whose inputs do not depend on memory values (e.g., they depend only on processor register values or values defined in the code itself), and (3) deterministic instructions whose inputs depend on values read from memory. Thus, in some embodiments, the execution of a reconstructed instruction can be accomplished using solutions to the following three corresponding challenges: (1) how to record non-deterministic instructions that produce outputs not completely determined by their inputs, (2) how to reproduce the values of the input registers for instructions that depend on registers, and (3) how to reproduce the values of the input memory for instructions that depend on memory reads.

[0037] As a solution to the first challenge of how to record "non-deterministic" instructions executed by a thread that do not produce completely predictable outputs (as their outputs are not completely determined by the data in general registers or memory), embodiments include storing the side effects of the execution of such instructions in the trace of the thread. As used herein, "non-deterministic" instructions include less common instructions that: (i) produce non-deterministic outputs each time they are executed (i.e., RDTSC on an Intel processor, which writes the number of processor cycles since the last processor reset into a register), (ii) can produce deterministic outputs, but depend on inputs not traced by the recording component 106a (e.g., debug registers, timers, etc.), and / or (iii) produce processor-specific information (e.g., CPUID on an Intel processor, which writes processor-specific data into a register). Storing the side effects of the execution of such instructions can include, for example, storing the register values and / or memory values changed by the execution of the instruction. In some architectures such as those from Intel, processor features (such as those found in virtual machine extensions (VMX)) can be used to capture instructions for recording their side effects in the trace file(s) 108.

[0038] The solution to the second challenge of reproducing the input register values for deterministic instructions executed by a thread (e.g., whose input depends only on processor register values) is straightforward because they are the output of the execution of (one or more) previous instructions in the thread. Recording the execution of the entire series of processor instructions in the thread's trace can thus be reduced to reproducing the register values at the start of the series; the (one or more) trace files 108 do not need to store a record of which specific instructions were executed in the series, or intermediate register values. This is because the actual instructions are available in the application code 107 itself and can be made available during replay. These instructions can thus be provided with the recorded inputs (i.e., the initial set of recorded register values) during replay and executed in the same manner as they were executed during the thread's trace.

[0039] As a solution to the third challenge of reproducing the input memory values for deterministic instructions executed by a thread whose input depends on memory values, embodiments include recording in the thread's trace the memory values consumed by the instructions in the thread (i.e., the values it reads), regardless of how the values read by the instructions are written to memory. In other words, some embodiments include recording only memory reads and not memory writes. For example, although values may be written to memory by the current thread, by another thread (including the kernel, e.g., as part of handling an interrupt), or by a hardware device (e.g., input / output hardware 105), it is the values read by the thread's instructions that are needed for a complete replay of the instructions of the thread that performed the read. This is because it is the values read by the thread (and not necessarily all the values written to memory) that indicate how the thread was executed. Although in some embodiments the value of each memory value read may be stored in the (one or more) trace files 108, other embodiments include optimizations such as prediction techniques that attempt to predict the appropriate value without having to record every read. For example, in some implementations, if the predicted value is the value actually read from memory, nothing needs to be recorded in the (one or more) trace files 108; however, if the predicted value does not match the value actually read, the value read is recorded in the (one or more) trace files 108. One such prediction technique is to predict that the next memory value read by the thread will be the same as the value previously read by the thread, as described below. Another prediction technique is to always predict that the next memory read will have a value of zero. Other example prediction techniques are also discussed later.

[0040] Additionally, since each thread in a thread is logged independently of each other, (one or more) trace files 108 do not need to record a strict ordering of every instruction executed across all threads (e.g., using a totally sequentially consistent recording model as discussed above). However, an approximation of the order in which instructions are executed can be useful for later debugging. Thus, to record an approximation of the order in which instructions are executed across threads, embodiments instead define or identify a "trace memory model" that has defined "sortable" and "non - sortable" events. Then, the recording component 106a records the execution sequence of "sortable" events that occur during thread execution. For example, embodiments can use a monotonically increasing number ("MIN") that ensures no duplicates to record the order in which sortable events occur across threads. Generally, the trace memory model should define how threads can interact via shared memory and their shared use of data in memory. The trace memory model used can be the memory model defined by the programming language (e.g., C++14) used to compile the application code 107, or some other memory model defined for tracing purposes (such as the memory model defined by the time - travel debugger 106).

[0041] As an example, a first trace memory model can treat only kernel calls, traps, and exceptions (from user mode) as sortable. This trace memory model will have low overhead because these operations are each relatively "expensive" in themselves, they can be traced in any case and provide a very coarse - grained overview of the ordering.

[0042] A second example trace memory model can treat full fences (i.e., operations with both acquire and release semantics) as sortable. Examples of such operations can include Intel's "lock" instruction, kernel calls, exceptions, and traps. This memory model will provide sufficient ordering for almost all cross - thread communication that occurs in a process when code uses "interlocked" type primitives for cross - thread communication (which is common in operations such as those from WINDOWS from Microsoft Corporation).

[0043] A third example trace memory model can treat all acquires and releases as sortable. This memory model can be suitable for processor - based ARM instruction sets because ARM does not treat most loads and stores as acquires or releases. On other architectures such as those from Intel (where most memory accesses are acquires or releases), this would be equivalent to sorting almost all memory accesses.

[0044] A fourth example trace memory model can treat all memory loads as sortable. This will provide strong ordering but may result in degraded performance compared to other example memory models.

[0045] The foregoing memory model has been presented only as an example, and one of ordinary skill in the art will recognize that, given the disclosure herein, a wide variety of memory models can be selected.

[0046] In view of the foregoing, Figure 2 An example flowchart of a method 200 for recording a replayable trace of the execution of a multi-threaded process is illustrated. As depicted, method 200 includes an action 201 of identifying a memory model. Action 201 can include identifying a trace memory model that defines one or more sortable events that are to be ordered across multiple threads of a multi-threaded process. For example, action 201 can include recording that component 106a identifies or defines a memory model, such as, by way of example only: treating kernel calls, traps, and exceptions as sortable; treating full fences as sortable; treating all acquires and releases as sortable; treating all memory loads as sortable; and so on, as discussed above.

[0047] Figure 2 Also depicted is that method 200 includes an action 202 of simultaneously executing multiple threads. Action 202 can include simultaneously executing multiple threads across one or more processing units of one or more processors while observing their execution using recording component 106a. For example, action 202 can include executing a first thread of a process on a first processing unit 102a of a processor while executing a second thread of the process on a second processing unit 102a of the same processor. As another example, action 202 can include executing two threads of a process on the same processing unit 102a of a processor that supports hyper-threading. As another example, action 202 can include executing two threads of a process on different processor units 102a of different processors. Combinations of the foregoing are also possible.

[0048] Figure 2 Also depicted is that method 200 includes an action 203 of independently recording a trace for each thread. Action 203 can include independently recording a separate replayable trace for each thread during the execution of the multiple threads. As will be clear from the examples hereinbelow, the recorded traces for each thread are independent of one another, except that they can include sortable events identified by MIN across threads. As illustrated, when they execute, action 203 can include sub-actions performed by recording component 106a for each thread.

[0049] Figure 2It is also depicted that action 203 includes action 204 of recording the initial thread state for each thread. For example, action 204 may include recording that component 106a stores the initial processor register values. Other initial thread states may include the thread environment block (TEB) and / or the process environment block (PEB) for the first thread of the process. Recording the TEB and / or PEB may later provide useful debugging information (e.g., thread local storage data). Other initial thread states may include the execution stack of the thread, especially in the case where recording component 106a is starting to record an already executed thread.

[0050] Figure 2 It is also depicted that action 203 may include action 205 of recording the side effects of non-deterministic instructions. Action 205 may include recording the side effects of at least one non-deterministic processor instruction executed by the thread. For example, recording component 106a may record the side effects of non-deterministic instructions by recording any changes made to register values by the instructions.

[0051] Figure 2 It is also depicted that action 203 includes action 206 of recording memory reads for each thread. Action 206 may include recording at least one memory read executed by at least one processor instruction, where the at least one processor instruction executed by the thread takes the memory as an input. For example, recording component 106a may record the (one or more) trace files 108 that can be used to reproduce the values read from memory during replay time. This may include recording each read, or applying one or more algorithms to predict reads to reduce the number of read entries that need to be recorded in the (one or more) trace files 108.

[0052] Figure 2 It is also depicted that action 203 includes action 207 of recording sortable events for each thread. Action 207 may include using a monotonically increasing number to record at least one sortable event executed by the thread, where the monotonically increasing number sorts the event among other sortable events across the multiple threads. For example, recording component 106a may record the execution sequence of events that can be sorted using MIN applied across threads by a trace memory model. Thus, replay component 106b can ensure that these events are ordered across threads during replay. An example of a sortable event is the start of a trace. Other examples are given in conjunction with Tables 1 and 2 below.

[0053] As illustrated by the double-headed arrows between actions 205, 206, and 207, these actions may occur in any order and may occur multiple times during the tracing of a thread, as will become clear in conjunction with the following examples.

[0054] When recording a trace, the recording component 106a may employ compression techniques to reduce the size of the trace file(s) 108. For example, the recording component 106a may dynamically compress the trace data before writing the trace data to a storage device (whether in the system memory 103 or in the data store 104). Alternatively, the recording component 106a may compress the trace file(s) 108 at the end of the recording.

[0055] Tables 1 - 5 illustrate specific examples of using the method 200 to trace the execution of a process. Specifically, Tables 1 - 5 illustrate example techniques for recording the execution of a single thread of a multi - threaded process. Although only one thread is recorded in this example for simplicity and clarity, the recording techniques do not need to change regardless of the number of threads that modify the program state. Additionally, even if there are external entities (e.g., the kernel, hardware) that modify the shared memory, the recording will not need to change. This is because these recording techniques record what the processing unit 102a actually sees (reads) and does when executing a thread, rather than focusing on constraining the execution to produce a completely predictable sequence of executed instructions.

[0056] First, Table 1 illustrates an example list of processor instructions (e.g., from the application code 107) to be executed by the processing unit 102a as part of a thread.

[0057] Table 1: List of Processor Instructions

[0058]

[0059]

[0060] In Table 1, the "Address" column refers to the memory address where the instruction specified in the "Instruction" column is found (e.g., in the instruction part of the cache 102b). In this example, these addresses are simplified to two digits. In the instructions, R1 - R7 refer to the processor registers 102c, and the data enclosed in square brackets ("[]") refers to a memory location (e.g., in the data part of the cache 102b).

[0061] Table 2 illustrates the sequence of instruction execution for the thread, including the register states, and sample data that may be recorded in the trace file(s) 108 to support the replay of the execution of the thread.

[0062] Table 2: Instruction Execution Sequence

[0063]

[0064]

[0065]

[0066] In Table 2, the "ID#" column refers to the specific sequence of instruction executions, and the "Address" column indicates the instruction address of the instruction executed at the ID# (see Table 1). Although not necessarily recorded in the trace file(s) 108, the columns "R1" - "R7" are included to illustrate the state changes to some of the processor registers as an aid to understanding program execution. The "Trace" column indicates the type of data that can be recorded in conjunction with instruction execution.

[0067] According to method 200, the recording component 106a will identify a trace memory model that defines "sortable" and "non - sortable" instructions / events (action 201). This enables the recording component 106a to record ordered (sortable) events across threads according to the MIN. Later, during replay, this enables the replay component 106b to ensure that these sortable instructions are replayed across threads in the proper order relative to each other, and also enables the replay component 106b to have flexibility in non - sortable instructions within and across replay threads.

[0068] For illustration, Figure 3 Example 300 of event ordering across simultaneously executing threads (Thread A and Thread B) is provided. Specifically, Figure 3 A timeline 301 of the execution of Thread A and a simultaneous timeline 302 of the execution of Thread B are illustrated. These timelines 301, 302 show the relative execution order of seven sortable events. Specifically, timeline 301 shows Thread A executing sortable events 3 and 5, and timeline 302 shows Thread B executing sortable events 1, 2, 4, 6, and 7. During replay, the replay component 106b ensures that these sortable events are replayed in the proper order within their respective threads and in the proper order across the two threads (i.e., B1 => B2 => A3 => B4 => A5 => B6 => B7).

[0069] The non-orderable instructions are replayed relative to these orderable instructions. For example, the replay component 106b will replay any non-orderable events in time block 303 of thread B before any non-orderable events in time block 304 of thread B. Additionally, the replay component 106b will replay the non-orderable events in time block 303 of thread B before any non-orderable events in time block 305 of thread A. Further, the replay component 106b will replay the non-orderable events in time block 305 of thread A before any non-orderable events in time block 306 of thread B. However, the replay component 106b cannot enforce any specific ordering between the replay of non-orderable events in time block 305 of thread B and the replay of non-orderable events in time block 304 of thread B because that ordering cannot be determined solely based on orderable events. However, the replay component 106b can enforce some ordering between these instructions when both of these instructions are going to access shared memory because their ordering can be determined at least in part based on how and when they access shared memory (including the values accessed).

[0070] Continuing to refer Figure 2 , method 200 will proceed to record an individual trace for each thread (action 203). Before executing any instructions for a given thread, the recording component 106a records an initial state / context (action 204), such as the initial values of the processor registers 102c (i.e., in this example, all 1 (0xff) for each of R1 - R7).

[0071] The recording component 106a then proceeds to monitor and record the execution of the thread, recording appropriate instructions / events in the (one or more) trace files 108 when they occur, such as side effects of the execution of non-deterministic instructions (action 205), memory reads (action 206), and orderable events (action 207).

[0072] For example, at ID #0, the recording component 106a records the execution of the instruction at address 10 (i.e., according to Table 1, "Move R1 <- {Timer}" (move R1 <- {timer})). This instruction reads the timer and places the value in register (R1). In this example, assume that the selected trace memory model (action 201) determines that the timer access is an "orderable" instruction that should be recorded by the MIN. Thus, the recording component 106a records the ordered (orderable) event "X" in the (one or more) trace files 108 (action 207). In this example, this is marked in the "Trace" column as "Instr:0 <Orderable event,X> (Instr:0 <Orderable event,X>)", meaning that zero instructions recorded since the last data written to the trace identify the ordered (orderable) event identified by MIN X.

[0073] Note that the specific symbol (MIN) used to record the ordered (orderable) event can be any symbol (hence the use of the generic "X"), as long as it is a value that is guaranteed not to repeat within the same record and is monotonically increasing (e.g., A, B, C, etc.; 1, 2, 3, etc.; 10, 20, 30, etc.). One possible option is to use the timer value provided by the processor, such as the value returned by the RDTSC instruction on an Intel processor. Thus, operations that occur after "X" will have a sequence identifier that is strictly greater than "X", regardless of which thread those operations occur on.

[0074] Additionally, instruction 10 is a first class of non-deterministic instructions whose result depends not only on the input (e.g., there is no input, and the timer value returned each time the instruction is executed will be different). Thus, the recording component 106a records the side effects of its execution in the trace (action 205). In this example, this is marked in the "Trace" column as "Instr:0 <Side effects,R1=T1,PC=11> (Instr:0 <Side effects,R1=T1,PC=11>)", meaning that zero instructions recorded since the last data written to the trace (which was "Instr:0 <Seq.event,X> (Instr:0 <Ordered event,X>)") record the new value of R1 (which is time T1), and increment and record the program counter (PC, not shown) register (i.e., to the next instruction at address 11).

[0075] Although this example records the update to the PC as a side effect, some embodiments may omit doing so because the replay component 106b will be able to determine how far the PC should be incremented based on the analysis of the executed instructions. However, recording the PC allows "jump" instructions to be recorded as side effects, which provides the opportunity to record the execution of several instructions using a single entry in the (one or more) trace files 108.

[0076] At this point, trace file(s) 108 have sufficient information so that replay component 106 b can establish an initial thread state (e.g., set the values of R1-R7 to 0xff), can replay instruction 10 (by reproducing its side effects), and can order instruction 10 relative to any other instructions within that thread or within other threads that are tracking the ordering of concurrently executing threads as recorded component 106 a tracks.

[0077] Next, at ID#1, instruction 11 performs a read (reads the address of memory location A and places the value in R2, see Table 1). Since this instruction reads memory, the recording component 106a records this to the trace file(s) 108 (action 206) in such a way that upon replay the replay component 106b can reproduce or predict the value that was read. As previously discussed, this can include the recording component 106a recording the value of each read, but there are various optimizations to reduce the amount of data that needs to be written. One approach is to predict that the value read by the current instruction is the value that was last read by the thread (and recorded in the trace file(s) 108). In this case, the trace file(s) 108 do not contain a previous read for the thread, so the read needs to be recorded. In the example, this is marked as "Instr:1<Read valueptrA> (Instr:1<read value ptrA>)", which means that since the last data written to the trace (i.e., "Instr:0<Sideeffects,R1=T1,PC=11> (Instr:0<side effect, R1=T1, PC=11>)") reads the address of memory location A. As indicated, this address (ptrA) will be placed into R2 by the instruction.

[0078] The instructions executed at ID#2 and ID#3 (instructions 12 and 13) also need to be recorded because they are also reads (action 206). The data recorded by the recording component 106a in the trace file(s) 108 is similar to the data recorded for instruction 11. For ID#2 (instruction 12), the tag "Instr:1<Read value ptrB> (Instr:1 <read value ptrB>)" means one instruction since the last data written to the trace, read the address of memory location B. For ID#3 (instruction 13), the mark "Instr:1<Read value 5> (Instr:1 <read value 5>)" means that one instruction since the last data written to the trace read the value 5. Note that according to the example prediction algorithm, both of these will need to be written to trace file(s) 108 because each read produces a different value than the previous read. As indicated, the address for location B (ptrB) will be placed into R2 by instruction 12, and the value 5 will be placed into R4 by instruction 13.

[0079] For ID#4 and ID#5, instructions 14 and 15 are deterministic, with no reads from memory, and therefore recording component 106a does not need to record anything in trace file(s) 108. Similarly, for ID#6, instruction 16 is deterministic and depends only on the values of R5 and R4. Recording in trace file(s) 108 is not necessary because these values will be reproduced by replay component 106b based on the replay of instructions 13 and 14 together with the trace data for the instruction at ID#3.

[0080] For ID #7, instruction 17 also does not require any trace data to be recorded, as its behavior (no jump is taken because the value of R5 is not less than the value of R4) is completely determined by data already recorded (or implied) in trace file(s) 108. For example, upon replay, the trace data for ID #3 will cause replay component 106b to place a value of 5 into R4, and a value of zero will have been written to R5 based on replaying instruction 14 at ID #4.

[0081] For ID#8, instruction 18 is a read from memory (the memory location identified by [R2+R5]) and should therefore be considered logged (action 206). Assume that the value 5 is observed by logging component 106a as having been read during execution (as indicated in column "R7"). Although logging component 106a may log this read in trace file(s) 108 (e.g., "Instr:5<Read value 5> (Instr: 5 <read value 5>)"), but it can avoid doing so by applying the read prediction algorithm discussed above: if the value read is the same as the value from the last read operation performed by the thread, then nothing is recorded. In this case, the last read value is 5 (ID #3), so for ID #8, the recording component 106a does not need to record anything in the trace file(s) 108.

[0082] For ID#9-ID#11, the recording component 106a also does not need to record anything to be able to replay instructions 19-21, because everything required to execute these instructions is already known. For ID#12 (instruction 21), the recording component 106a also does not need to record anything, because the inputs to this instruction are only registers, and because writes to memory (i.e., the location identified by [R3+R6]) do not need to be recorded (as previously discussed), so although reads from memory affect how and which instructions are executed, writes do not. For ID#12-ID#16 (instructions 23, 24, Loop: 16, 17), the recording component 106a also does not need to record anything, because everything required to execute these instructions is already known.

[0083] For ID#17, the second time instruction 18 is encountered (a read from memory), a value of 0 is read (as indicated in column "R7"). This value is not what would be predicted by the logging component 106a using the sample prediction algorithm (because it is different from the last value read, 5), so the logging component 106a adds it to the trace file(s) 108 (e.g., "Instr:14").<Read value 0> (Instr: 14 <read value 0>)", which indicates the fourteenth instruction after the last data entered, read value zero).

[0084] For ID#18 - ID#22 (Instruction 18, 20, 21, Loop:16 (Loop: 16), and 17), the recording component 106a does not need to add anything to the (one or more) trace files 108. As before, the replay engine 106b will already have enough information to reproduce the same results observed by the recording component 106a. During replay, the replay component 106b compares the same values, takes the same jumps, and so on.

[0085] ID#23 is the third encounter of Instruction 18 (read from memory), and another 0 is read (as indicated in column "R7"). Although the recording component 106a needs to consider adding it to the (one or more) trace files 108 (Action 206), the value is predicted by the prediction algorithm (the last read at ID#17 was also 0), so the recording component 106a does not record anything into the (one or more) trace files 108.

[0086] For ID#24 - ID#28 (Instruction 19, 20, 21, Loop:16 (Loop: 16), and 17), the recording component 106a also does not need to add anything to the (one or more) trace files 108 because the replay component 106b will already have enough information to reproduce the same results.

[0087] For ID#29, the fourth encounter of Instruction 18 (read from memory), a value of 2 is read (as indicated in column "R7"). This value is not the predicted value (it is different from the last value read, which was 0), so the recording component 106a adds it to the (one or more) trace files 108 according to Action 206 (twelve instructions since the last entry, read value 2).

[0088] For ID#30 - ID#37 (Instruction 19, 20, 21, 22, 23, 24, Loop:16 (Loop: 16), and 17), the recording component 106a also does not need to add anything to the (one or more) trace files 108. Again, the replay component 106b already has enough information to reproduce the same results.

[0089] ID#38 (Instruction 18) is another memory read (from location [R2 + R5]). As indicated in R7, a value of 2 is read, which is the same as the last read at ID#29. Although the recording component 106a needs to consider writing it to the (one or more) trace files 108 (Action 206), according to the prediction algorithm it does not need to do so.

[0090] For ID#39 - ID#46 (Instructions 19, 20, 21, 22, 23, 24, Loop:16 (Loop: 16), and 17), the recording component 106a also does not need to add anything to the (one or more) trace files 108. Again, the replay component 106b already has enough information to reproduce the same results observed by the recording component 106a.

[0091] At ID#9, the recording component 106a observes another timer read at instruction 26. Since this is a "sortable" event according to the memory model (as discussed in conjunction with ID#0), the recording component 106a uses an incrementing identifier according to action 207 (e.g., stating the data of the eighteenth instruction after the data logged last, the ordering event X+N that occurred) to record the sequence / order of operations. Additionally, since this is a non-deterministic instruction, the recording component 106a records its side effects according to action 205 (e.g., stating the data of zero instructions after the data logged last, recording effect: R2 = T2 and PC = 26).

[0092] For ID#48, there is no read and the instruction is deterministic, so the recording component does not log anything in the (one or more) trace files 108.

[0093] Table 3 outlines examples of trace data that the recording component 106a might have recorded in the (one or more) trace files 108 as part of tracing the execution of this thread.

[0094] Table 3: Example Traces with Predictions

[0095] <Initial context(e.g.,registers)>(<Initial context (e.g., registers)>) Instr:0<Orderable event,Id X>(Instr:0<Orderable event,Id X>) Instr:0<Side effects,R1=T1,PC=11>(Instr:0<Side effects,R1=T1,PC=11>) Instr:1<Read value ptrA>(Instr:1<Read value ptrA>) Instr:1<Read value ptrB>(Instr:1<Read value ptrB>) Instr:1<Read value 5>(Instr:1<Read value 5>) Instr:14<Read value 0>(Instr:14<读取值0>) Instr:12<Read value 2>(Instr:12<Read value 2>) Instr:18<Orderable event,Id:X+N>(Instr:18<Orderable event,Id:X+N>) Instr:0<Side effects,R2=T2,PC=26>(Instr:0<Side effects,R2=T2,PC=26>)

[0096] The replay component 106b can later use this trace data to establish an initial context (e.g., register state) and then execute the same instructions with the thread's code (e.g., Table 1) in the same way they were executed at recording time by reproducing the side effects of the nondeterministic instructions and providing memory values as needed (using knowledge of the prediction algorithms used during recording). Additionally, the ordering events (X and X+N) enable the instructions at ID#0 and 47 to be executed in the same order relative to ordering events in other threads as they were executed during recording. Thus, the tracing mechanism supports tracing and replay of individual executing threads and multiple simultaneously executing threads while recording a small amount of trace data and much less than a full deterministic recording of the executed instructions and full memory state. In particular, in the foregoing example, the recording component 106a does not track or record when, why, or who (i.e., another thread in the process, the same thread earlier, another process, the kernel, the hardware, etc.) writes the values read and consumed by the code in Table 1. However, the replay component 106a is still enabled by the trace table 3 to replay the execution in the exact manner it was observed.

[0097] As will be appreciated by those skilled in the art, given the disclosure herein, there are many variations on the specific ways in which trace data can be recorded. For example, although the example trace tracks when events occur based on instruction counts relative to the number of instructions executed since the previous entry in the trace, absolute instruction counts (e.g., a count starting at the beginning of the trace) can also be used. Other examples that can be used to uniquely identify each instruction executed can be based on a count of the number of CPU cycles executed (relative to the previous entry, or as an absolute value), a count of the number of memory accesses made, and / or a jump count taken in conjunction with the processor's program counter (which can be relative to the last non-sequential or kernel call, or absolute from a defined time). When using one or more of the foregoing techniques, care may be needed when recording certain types of instructions (such as "repeat" instructions (e.g., REP on Intel architectures)). Specifically, on some processors, repeat instructions can actually execute with the same program counter on each iteration, and thus on these processors, a repeat instruction with multiple iterations can only count as one instruction. In these cases, the trace should include information that can be used to distinguish each iteration.

[0098] Additionally, there are various mechanisms for recording and predicting the values read. As a simple example, an embodiment could eliminate the prediction algorithm entirely. Although this would result in a longer trace, it would also remove the limitation that the replay component 106b is configured with the same prediction algorithm as the recording component 106a. Table 4 illustrates what a trace might look like without using a prediction algorithm, and where each memory read is recorded. As shown, without a prediction algorithm, the trace would include three new entries (highlighted), and the other entries (i.e., those after each new entry) each have an update to the instruction count that reflects the count from the new entry in front of it.

[0099] Table 4: Example Trace Without Prediction

[0100] <Initial execution context(e.g.,registers)>(<Initial execution context (e.g., registers)>) Instr:0<Orderable event,Id X>(Instr:0<Orderable event,Id X>) Instr:0<Side effects,R1=T1,PC=11>(Instr:0<副作用,R1=T1,PC=11>) Instr:1<Read value ptrA>(Instr:1<Read value ptrA>) Instr:1<Read value ptrB>(Instr:1<Read value ptrB>) Instr:1<Read value 5>(Instr:1<Read value 5>) Instr:5<Read value 5>(Instr:5<Read value 5>) Instr:9<Read value 0>(Instr:9<读取值0>) Instr:6<Read value 0>(Instr:6<Read value 0>) Instr:6<Read value 2>(Instr:6<Read value 2>) Instr:9<Read value 2>(Instr:9<Read value 2>) Instr:9<Orderable event,Id:X+N>(Instr:9<Orderable event,Id:X+N>) Instr:0<Side effects,R2=T2,PC=26>(Instr:0<副作用,R2=T2,PC=26>)

[0101] Note that the trace file(s) 108 need not record the memory addresses of the values read if they can be mapped based on which instruction consumed them. Thus, an instruction that generates more than one read may need a means to identify which read is the one in the trace file(s) 108 (i.e., when there is only one read for that instruction) or which read is which (i.e., in the case where there are several reads in the trace file(s) 108). Alternatively, the trace file(s) 108 can contain the memory addresses for these reads or for all reads. The trace file(s) 108 only need to include enough information such that the replay component 106b can read the same values that the recording component 106a observed and match them to the same parameters so that it can produce the same result.

[0102] Additionally, events that are not discoverable in the code flow (i.e., discontinuities in code execution) can occur, such as access violations, traps, interrupts, calls to the kernel, etc. The recording component 106a also needs to record these events so that they can be replayed. For example, Table 5 illustrates an example trace that could be recorded in the trace file(s) 108 in the case where an access violation has occurred at ID#33 (instruction 22).

[0103] Table 5: Example Trace with Exception

[0104]

[0105]

[0106] Specifically, the trace contains: "Instr:4 <Exception record> (Instr:4 <Exception record>)" which indicates that an exception occurred and the exception occurred at the 4th instruction after the last entry in the trace; and "Instr:0 <exceptioncontext>(Instr:0 <Abnormal context>)", which, in combination with restarting execution after an abnormality and recording any appropriate state, resets the instruction count to zero. Although Table 5 shows separate records for indicating the occurrence of an abnormality and for recording the abnormal context for clarity, they can be in the same entry. Setting the instruction count to zero indicates to the replay component 106b that the two entries apply to the same instruction. Now, since the (one or more) trace files 108 contain entries for the abnormality, the trace has the exact location for such an abnormality. This enables the replay component 106b to raise an abnormality at the same execution point of the abnormality as observed during recording. This is important because abnormalities cannot be inferred by looking at the code flow (since their occurrence is often based on data not present in the code flow).

[0107] As mentioned above, various mechanisms for tracking the values read can be employed by the recording component 106a, such as predicting that the value to be read is equal to the value read last time, and recording the read in the (one or more) trace files 108 if the values are different. An alternative approach uses a cache to extend the prediction algorithm such that the recording component 106a predicts that the most likely value to be read from a particular address is the last value read from or written to that address. Thus, this approach requires maintaining a cache of the memory range of the process and using the cache to track memory reads and writes.

[0108] To avoid the overhead of this approach (i.e., maintaining a cache of the last values placed in the entire memory range of the program), an improvement is that the recording component 106a creates a "shadow copy" of a much smaller amount of memory than the full memory being addressed (as discussed below). Then, for each read observed, the recording component 106a compares the value read by the instruction with the value at the matching location in the shadow copy. If the values are the same, there is no need for the recording component 106a to save a record of the read in the (one or more) trace files 108. If the values are different, the recording component 106a records the value in the (one or more) trace files 108. The recording component 106a can update the shadow copy on reads and / or writes to increase the likelihood of correctly predicting the value on the next read of the memory location.

[0109] In some embodiments, the size of the shadow copy can be limited to 2^N addresses. Then, to match an address with its shadow copy, the recording component 106a obtains the low N bits of the memory address (i.e., where N is a power of 2, which determines the size of the shadow copy). For example, in the case of a value of N=16, the shadow copy will be 64k (2^16) addresses, and the recording component 106a obtains the low 16 bits of each memory address and compares them to the offset in the shadow copy. In some embodiments, the shadow copy is initialized with all zeros. Zero can be selected because it is an unambiguous value and because zero is read from memory quite frequently. However, other choices of initial values can alternatively be used depending on the implementation.

[0110] Figure 4 Illustrated is an example 400 of using shadow copies according to some embodiments. Figure 4 4 (sixteen) memory locations, which is only a fraction of the locations that a typical process will be able to address on a contemporary computer (which will typically be on the order of 2^32, 2^64, or more memory locations). In the representation of addressable memory 401, the address column specifies the binary address of each memory location (i.e., binary addresses 0000-1111), and the value column indicates the location for the data to be stored at each memory location. Based on the shadow copy described above, Figure 4 Also shown is a corresponding shadow copy 402 of addressable memory 401 that can be maintained by recording component 106a during tracing of a thread. In the example, the value N is equal to two, so shadow copy 402 stores 2^2 (four) memory locations (ie, binary addresses 00-11).

[0111] Whenever the logging component 106a detects a read from the addressable memory 401, the logging component 106a compares the value read from the addressable memory 401 with the corresponding location in the shadow copy 402. For example, when any of the memory locations 401a, 401b, 401c, or 401d (i.e., binary addresses 0000, 0100, 1000, or 1100) is read, the logging component 106a compares the value read from that location with the value in location 402a of the shadow copy (i.e., binary address 00, because the last N digits of each of the aforementioned memory addresses are 00). If the values match, the read does not need to be recorded in the trace file(s) 108. If the values do not match, the logging component 106a records the read in the trace file(s) 108 and updates the shadow copy with the new value.

[0112] Even though each location in the shadow copy 402 represents multiple locations (in this case, four) in the addressable memory 402, it should be noted that most programs may perform multiple reads from the same location in the addressable memory 402 and / or multiple writes to the same location in the addressable memory 402 (e.g., to read or update a variable value), and thus there may be a large number of reads predicted by the shadow copy 402. In some embodiments, the recording component 106a may also update the shadow copy 402 when a write occurs, which can further increase the likelihood of correct prediction.

[0113] In some embodiments, rather than tracking memory reads at a fine-grained level of memory addresses, the recording component 106a may track reads across multiple threads at the memory page level (e.g., based on a page table). This embodiment is based on the recognition that memory pages are restricted such that each page of memory can (i) be written by one thread without any other thread having read or write access to that page; or (ii) be read by as many threads as needed, but no thread can write to it. Embodiments can thus group threads into families such that during the recording period, threads within a family always execute non-concurrently with respect to each other (i.e., two threads of the same family cannot execute concurrently). Then, the above-mentioned restrictions are applied across thread families such that, except for concurrent page accesses, threads of different thread families can run concurrently with respect to each other. If a thread family "owns" a page for writing, no other thread family can have access to it; however, one or more thread families can share a page for reading.

[0114] When a thread family accesses a page for reading or writing, it is necessary to know the overall content of the page. If the page was produced / written by a thread that has already been recorded in the recording component 106a, the recording component 106a already knows the content of the page. Thus, in some embodiments, the recording component 106a places only information identifying the point at which the writing thread in the record releases the page on the trace. For pages that have been produced / written by an external entity (e.g., the kernel, an untracked thread, a hardware component, etc.), the strategy for recording the page such that it is available during replay can include the recording component 106a recording the entire page or the recording component 106a recording a compressed version of the page. If the page has been previously recorded, another strategy includes the recording component 106a storing only the difference in the page values between the current version of the page and the previously recorded version.

[0115] As previously indicated, when debugging code traced by the recording component 106a, breakpoints in the "backward" direction are replayed by the replay component 106b from the time before the breakpoint until the (one or more) trace files 108 reach the breakpoint (e.g., the last time the breakpoint was hit was before the debugger is currently analyzing the code flow). Using the tracing described above, this would mean replaying the trace from the start of the (one or more) trace files 108. Although this is acceptable for smaller traces, it can be time-consuming and inefficient for larger traces. To improve the performance of the replay, in some embodiments, the recording component 106a records multiple "keyframes" in the trace file 108. The keyframes are then used by the replay component 106b to more granularly "zero in" on the breakpoint. For example, in some implementations, the replay component 106b can iteratively "rewind" increasing numbered keyframes (e.g., doubling the number of the keyframe each iteration) until the selected breakpoint is reached. By way of illustration, the replay component can rewind one keyframe and attempt to hit the breakpoint, if it fails it can rewind two keyframes and attempt to hit the breakpoint, if it fails it can rewind four keyframes and attempt to hit the breakpoint, and so on.

[0116] Generally, a keyframe includes enough information for the replay component 106b to replay the execution starting at the keyframe, generally not caring what occurred in the trace before the keyframe. The exact timing of recording the keyframe and the data collected can vary based on the implementation. For example, at periodic intervals (e.g., based on the number of instructions executed, based on processor cycles, based on elapsed clock time, based on the occurrence of "sortable" events according to the trace memory model, etc.), the recording component 106a can record enough information for a keyframe in the (one or more) trace files 108 for each thread to replay the trace of each thread starting from the keyframe. This information can include the state of the hardware registers when the keyframe was acquired. This information can also include any information needed to place the memory prediction strategy in a known state so that reads starting at the keyframe can be reproduced (e.g., by recording the (one or more) memory snapshots, (one or more) shadow copies, (one or more) last read values, etc.). In some embodiments, keyframes can be used to support gaps in the trace. For example, if the trace is stopped for any reason, this can be marked in the (one or more) trace files 108 (e.g., by inserting a note that the trace was stopped and appropriately formed keyframes), and then the trace can be restarted later from that point in the (one or more) trace files 108.

[0117] In some embodiments, some key frames may include information that may not be strictly necessary to support replay at the key frame, but which proves useful during debugging (e.g., to help the time travel debugger 106 consume data generated during replay and present it in a useful form). This information can include, for example, copies of one or more portions of the stack of a thread, copies of one or more portions of memory, the TEB of a thread, the PEB of a process, the identities of loaded modules and their headers, etc. In some embodiments, key frames that include this additional information (e.g., "complete" key frames) may be generated and stored less frequently than conventional key frames (e.g., "lightweight" key frames). The frequency of collection of lightweight key frames and / or complete key frames, as well as the specific data collected in each, may be user-configurable at recording time.

[0118] In some embodiments, recording "complete" key frames may also have features for things such as reusable circular buffers, which are discussed below in connection with Figure 5 Recording "lightweight" key frames also enables the replay component 106b to replay in parallel; each thread trace after a key frame can be independent of the other threads and thus replayed in parallel. Additionally, key frames can enable the replay component 106b to replay different segments of the same thread trace in parallel. For example, doing so can be useful for using the time travel debugger 106 to hit breakpoints more quickly. In some embodiments, the recording of "complete" key frames is coordinated across threads (i.e., the recording component 106a records complete key frames for each thread of a process at approximately the same time), while "lightweight" key frames are recorded independently for each thread (i.e., the recording component 106a records lightweight key frames for each thread that are convenient or otherwise meaningful for that thread). Adjusting the conditions for recording different types of key frames provides flexibility for balancing trace size, replayability, recording performance, and replay performance.

[0119] Some embodiments may include the use of "snapshots", which include a complete copy of the relevant memory of a process. For example, the recording component 106a may obtain an initial snapshot of the memory of a process when starting a trace of that process. This enables the replay component 106b to provide the user of the time travel debugger 106 with the values of all memory locations used by the process, not just those observed to be accessed during recording.

[0120] In some embodiments, the (one or more) trace files include information that can be used by the replay component 106b to verify that the state of the program at replay time indeed matches the program state that existed during recording. For example, the recording component 106a may include periodic information in the (one or more) trace files 108, such as copied register data and / or calculated hashes of register data. This information may be included in key frames, may be included periodically along with standard trace entries, and / or may be included based on the number of executed instructions (e.g., every X instructions are placed in the trace), etc. During replay, the replay component 106b may compare the recorded register data (and / or calculated hashes) at the corresponding point in execution with the state data generated during replay to ensure that the execution states are the same (i.e., if the register data and / or calculated hashes of the data during replay match, the execution states are likely the same; if they do not match, the execution states have deviated).

[0121] As previously mentioned, some embodiments include recording the (one or more) trace files that include a "circular buffer" of limited capacity. Specifically, the circular buffer only records the tail portion of program execution (e.g., the last N minutes, hours, days, etc. of execution). Conceptually, the circular buffer adds new trace data to the front / top of the trace and removes old trace data from the back / bottom of the trace. For example, some applications may run for days, weeks, or even months before a program defect manifests. In such cases, tracking the complete history of program execution may be impractical (i.e., in terms of the amount of disk space used) and unnecessary, even with the compactness of the trace files recorded by the disclosed embodiments. Additionally, the use of a circular buffer may potentially allow the (one or more) overall trace files 108 to be stored in RAM, which can greatly reduce disk I / O and improve performance (both during recording and replay).

[0122] When implementing a circular buffer, embodiments may track "permanent" trace information and "temporary" trace information, where the "temporary" information is stored in the circular buffer. Examples of "permanent" trace information may include general information, such as the identity of loaded modules, the identity of the process being recorded, etc.

[0123] Figure 5 Example 500 illustrates a reusable circular buffer 501 according to one or more embodiments. In the example, each blank rectangle in circular buffer 501 represents a standard trace entry in the circular buffer 501 (e.g., such as an entry in Table 3 above). Each shaded rectangle (e.g., 504a, 504b, 504c, 504d, 504e) represents a key frame. The frequency of the key frames and the number of entries between key frames will vary based on the implementation. As indicated by arrow 502 and the dashed entries it covers, new entries (both standard entries and key frames) are added to one end of the buffer. As indicated by arrow 503 and the dashed entries it covers, the oldest existing entry (both standard entries and key frames) is removed from the other end of the buffer. Entries can be added / removed on a one-to-one basis or in chunks (e.g., based on elapsed time, number of entries, occurrence of key frames, etc.). The overall size of the buffer can be configured based on the desired length of the tail period of program execution. To replay from circular buffer 501, replay component 106b initializes the state data with the desired key frames and then replays the program execution from them.

[0124] As previously described, the key frames of circular buffer 501 can include "full" key frames, which not only enable replay component 106b to replay from each key frame (e.g., using the register values stored in the key frame), but also to use additional debugging features from each key frame (e.g., using additional information such as the TEB of the thread, data cache, etc.). However, exclusively or in addition to "full" key frames, the key frames of circular buffer 501 can also include "lightweight" key frames.

[0125] The number of circular buffers used during recording can vary. For example, some embodiments can use separate circular buffers per thread (i.e., each thread is allocated multiple memory pages for recording trace data), while other embodiments can use a single circular buffer to record multiple threads. When using a single circular buffer, the traces of each thread are still recorded separately, but the thread records to a shared pool of memory pages.

[0126] In a second embodiment (using a single circular buffer to track multiple threads), each thread can obtain a page from a pool of pages allocated to the circular buffer and begin filling it with trace data. Once a thread's page is full, the thread can allocate another page from the pool. When adding key frames, some embodiments attempt to add them for each thread in the thread at approximately the same time. The key frames can then be used to assign a "generation" to the records. For example, data recorded before the first key frame can be "first generation" records, and data recorded between the first key frame and the second key frame can be "second generation" records. When the pages in the pool have been exhausted, the pages associated with the oldest generation of records (e.g., generation 1) can be released for reuse in future records.

[0127] Although the foregoing disclosure has primarily focused on recording information that can be used to replay the execution of code into one or more trace files 108, there are many other types of data that can assist the recording component 106a in writing to the one or more trace files. Examples have been given of writing key frames and additional debug information to the trace. Other types of information that the trace can be marked with can include timing information, performance counters (e.g., cache misses, branch misprediction, etc.), and the recording of events that do not directly affect the replay of the trace but can help with synchronization (e.g., embedding data that captures when the user interface was captured so this can be reproduced during replay). Additionally, when user mode code is being traced, the recording component 106a can use information such as the following to mark the one or more trace files 108: (i) when a user mode thread is scheduled in or out, (ii) when a user mode thread is suspended, (iii) when the user switches to focus on the application being traced, (iv) the recording of messages received by the application, (v) when a user mode process causes a page fault, and so on. Those of ordinary skill in the art will realize that, given the disclosure herein, the specific manner for recording any of the foregoing can vary based on the implementation.

[0128] The time machine debugger 106 can be implemented as a software component in various forms. For example, at least one or more parts of the time machine debugger 106 (e.g., the recording component) can be implemented as a code part injected into the runtime memory of the process being recorded (i.e., "instrumenting" the process being recorded), implemented as an operating system kernel component, implemented as part of a full machine emulator (e.g., BOCHS, Quick Emulator (QEMU), etc.), and / or implemented as part of a hypervisor (e.g., HYPER-V from Microsoft, XEN on LINUX, etc.). When implemented as part of an emulator or hypervisor, the time machine debugger 106 can be enabled to track the execution of the entire operating system. Thus, the time machine debugger 106 can track the execution of user mode code (e.g., when implemented as part of injected code or the kernel), and / or track the execution of kernel mode code and even the entire operating system (e.g., when implemented as part of a hypervisor or emulator).

[0129] Implementation Based on Processor Cache

[0130] Although the time machine debugger 106 can be implemented entirely in software, some embodiments include a hardware cache-based recording model, which can further reduce the overhead associated with recording the execution of a program. As before, this model is based on the general principle that the recording component 106a needs to create a recording (i.e., the (one or more) trace files 108) that enables the replay component 106b to replay each instruction executed during recording in the proper order and in such a way that each instruction produces the same output as it produced during recording. As is clear from the above disclosure, an important component of the (one or more) trace files 108 includes data that can be used to reproduce memory reads during replay. As discussed above, embodiments of recording such data can include recording each value read, using a prediction algorithm to anticipate the value read, using a shadow copy of the memory, recording memory pages and / or page table entries, etc.

[0131] Embodiments of the hardware cache-based recording model for tracking (including recording / replaying memory reads) are based on the observation that the processor 102 (including the cache 102b) forms a semi-closed or quasi-closed system. To further illustrate, Figure 6 An example computer architecture 600 for processor cache-based tracking is illustrated, which includes a processor 601 and a system memory 608, which can be mapped to Figure 1 the (one or more) processors 102 and system memory 103.

[0132] At a conceptual level, after the data is loaded into cache 603, the processor 601 can operate on its own for bursts of time as a semi-closed or quasi-closed system without any input. Specifically, during program execution, each processing unit 602 uses the data stored in the data cache 605 portion of cache 603 and executes instructions from the code cache 604 segment of cache 603 using register 607. For example, Figure 6 The figure shows that the code cache 604 includes a plurality of storage locations 604a - 604n (e.g., cache lines) for storing program code, and the data cache 605 includes a plurality of storage locations 605a - 605n (e.g., cache lines) for storing data. As previously discussed in conjunction with Figure 1 discussed, the cache can include multiple levels (e.g., level 1, level 2, and level 3, etc.). Although the cache 605 is depicted within the processor 601 for simplicity, it will be recognized that one or more portions of the cache 605 (e.g., level 3 cache) can actually exist outside of the processor 601.

[0133] When a processing unit 602 requires some influx of information (e.g., because it needs code or data that is not yet in cache 603), a "cache miss" occurs and the information is brought into cache 603 from an appropriate storage location in the system memory 608 (e.g., 608a - 608n). For example, if a data cache miss occurs when an instruction executes a memory operation at a memory address corresponding to location 608a (containing program runtime data), then that memory (i.e., the address and the data stored at that address) is brought into one of the storage locations of the data cache 605 (e.g., location 605a). In another example, if a code cache miss occurs when an instruction executes a memory operation at a memory address corresponding to location 608b (containing program code), then that memory (i.e., the address and the data stored at that address) is brought into one of the storage locations of the code cache 604 (e.g., location 604a). When new data is imported into cache 603, it can replace the information that was already in cache 603. In this case, the old information is pushed back to its proper address in the system memory 608. The processing unit 602 then uses the new information in cache 603 to continue execution until another cache miss occurs and new information is brought into cache 603 again.

[0134] Accordingly, embodiments of the hardware cache-based model for recording / playing back memory reads rely on the recognition that, apart from accesses to uncached memory (e.g., reads to hardware components and non-cacheable memory, as discussed later), all memory accesses performed by a process / thread during execution are performed through the processor cache 603. As a result, rather than creating a record that can play back each individual memory read as described above (e.g., a record of each read, a shadow copy, etc.), the recording component 106a can instead record (i) the data brought into the cache 603 (i.e., the memory address and the data stored at that address), and (ii) any reads to uncached memory.

[0135] Since the processor 102 can be regarded as a semi-closed or quasi-closed system, the execution flow within the processor that occurs during recording can be replicated by the replay component 106b using a simulator that emulates the processor and its cache system. The simulator is configured to produce the same results during replay as occurred during recording when the simulator is given the same inputs that occurred during recording by the replay component 106b. The simulator does not need to be an exact CPU simulator, as long as it emulates instruction execution (i.e., their side effects), and it does not need to match the timing, pipeline behavior, etc. of the physical processor. Thus, recording program execution can be streamlined to record: (i) a copy of the influx of data into the system (e.g., data brought into the cache 603 based on cache misses and uncached reads), (ii) how / when to apply the data to each input by the replay component 106b at the correct time (e.g., using the count of executed instructions), and (iii) data describing the system to be emulated (i.e., the processor 601, including its cache 603). This enables the time travel debugger 106 to model the processor 102 as a linear execution machine when recording a trace without having to record the internal parallelization or pipelining of instruction execution within the processor 102, and without having to maintain a record of the specific timing of the execution of each instruction within the processor 102.

[0136] As an overall overview, similar to the techniques described above, a hardware cache-based model for tracking threads begins by saving processor register values to a trace. Additionally, the model begins by ensuring that a suitable portion of the processor cache 603 for the thread is empty. As discussed below, the recording component 106a records data imported into the data cache 605 and can also record data imported into the code cache 604. Thus, the data cache 605 needs to be cleared at the start of the trace, and the code cache 604 needs to be cleared only if imports to the code cache 604 are being recorded. Then, at least a portion of the code of the thread is brought into the code cache 604 (e.g., by the processor performing a "cache miss" based on the memory address of the requested code, and if imports to the code cache 604 are being recorded, storing the imported cache line into the (one or more) trace files 108), and the processing unit 602 begins executing the processor instructions in the code cache 604.

[0137] When the code performs its first memory operation (e.g., a read or write to a memory address in the system memory 608), a "cache miss" occurs because the data cache 605 is empty and thus does not contain a copy of the memory at the accessed address. Accordingly, the correct portion of the system memory 608 is brought in and stored into a cache line in the data cache 605. For example, if the memory operation is addressed to a memory address at location 608a in the system memory 608, the data at the memory location 608a is written into a line in the data cache (e.g., location 605a). Since the data has been brought into the cache 603, the recording component 106a records the data (i.e., the address and the addressed data) into the (one or more) trace files 108.

[0138] After this point, future memory accesses bring new data into the cache (and thus are recorded by the recording component 106a into the (one or more) trace files 108) or are performed on data that has already been brought into the cache (e.g., they can read from or write to a cache line in the data cache 605). Subsequent reads of data that is already in the cache do not need to be recorded. Similar to the techniques described above in conjunction with Figures 1-5 the recording component 106a does not need to track writes to data in the cache because these writes can be reproduced by executing instructions with the recorded initial state and by reproducing the side effects of non-deterministic instructions.

[0139] Assume that the replay component 106b has access to the original code of the thread and that the execution was not interrupted during recording (e.g., there were no exceptions). Then, if the replay component 106b starts with an empty cache and the recorded register values, the replay component 106b can simulate the execution performed by the processor 601, including bringing the appropriate data into the cache 603 at the appropriate times and reproducing the side effects of non-deterministic instructions at the appropriate times.

[0140] Thus, like the techniques discussed above in conjunction with Figures 1-5 When using the cache-based model for recording, the recording component 106a still records the side effects of non-deterministic instructions in the (one or more) trace files 108 that do not have an output exclusively determined by their input values. Additionally, in some embodiments, the recording component 106a still selects the trace memory model and records the sortable events executed by the thread in the (one or more) trace files 108 with a monotonically increasing number that sorts these events across threads.

[0141] Additionally, as discussed in more detail later, in some embodiments, the recording component 106a tracks changes in the control flow of the thread that cannot be determined solely by the code of the thread. For example, a change in the control flow may occur due to an interrupt. These can include, for example, asynchronous procedure calls ("APCs"), calls from kernel mode, etc.

[0142] In addition, in some embodiments, the recording component 106a only tracks data cache 605 misses, while in other embodiments, the recording component 106a tracks both data cache 605 misses and code cache 604 misses. Specifically, if the thread is executing only non-dynamic code, only the cache lines imported into the data cache 605 need to be tracked because all the code to be executed is available during replay. However, if the thread includes dynamic code support, the cache lines imported into the code cache 605 also need to be tracked.

[0143] Furthermore, in some embodiments, the recording component 106a tracks which instructions perform memory reads from memory that is never cached / uncacheable because there is no cache miss to track these reads. Examples of uncached / uncacheable reads are reads from a hardware device or reads from memory that the processor 601 and / or the operating system otherwise consider uncacheable.

[0144] Figure 7 The flowchart of an example method 700 for using cached data to record the execution of executable entities for replayable tracing is illustrated. As depicted, method 700 includes an action 701 of simultaneously executing one or more threads. Action 701 may include executing one or more threads (e.g., user-mode threads, kernel threads, etc.) of an executable entity (e.g., a process, a kernel, a hypervisor, etc.) across one or more processing units of one or more processors, while observing their execution using a recording component 106a. For example, action 701 may include executing a first thread on a first processing unit 602 of a processor 601, while possibly executing a second thread on a second processing unit 602 of the same processor 601. As another example, action 701 may include executing two threads on the same processing unit 602 of a processor 601 that supports hyper-threading. As another example, action 701 may include executing two threads on different processing units 602 of different processors 601. The foregoing combinations are also possible.

[0145] Figure 7 Also depicted is that method 700 includes an action 702 of independently recording a separate trace for each thread. Action 702 may include: during the execution of one or more threads, independently recording a separate replayable trace for each thread. Thus, as illustrated, when they execute, action 702 may include sub-actions performed by the recording component 106a for each thread.

[0146] Figure 7 Also depicted is that action 702 includes an action 703 of recording an initial state for each thread. Action 703 may include recording the initial processor register state for the thread. For example, the recording component 106a may record the state of the register 607 corresponding to the processing unit 602 on which the thread executes. The initial state may also include additional information, such as a snapshot of the memory, stack information, the TEB of the thread, the PEB of the process, etc.

[0147] Figure 7 Also depicted is that operation 702 includes operation 704 of recording imported lines of cache data for each thread. Operation 704 may include: when a processor data cache miss is detected based on thread-based execution, recording at least one line of cache data imported into the processor data cache in response to the processor data cache miss. For example, during the execution of instructions in code cache 604, one of those instructions may perform a memory operation on a specific address in system memory 103 that is not yet in data cache 605. Thus, a "cache miss" occurs and the data at that address is imported into data cache 605 by processor 601. Recording component 106a creates a record of this data in trace file(s) 108. For example, if a memory address at location 608a is imported into cache line 605a, recording component 106a records the identity of the memory address and the data it contains in trace file(s) 108. If a thread executes dynamic code, a cache miss may also occur with respect to loading instructions into the processor's code cache 604. Thus, operation 704 may also include: recording at least one line of cache data imported into processor code cache 604 in response to a processor code cache miss.

[0148] Figure 7 Also depicted is that operation 702 may also include operation 705 of recording un-cached reads for each thread. Operation 705 may include: recording at least one un-cached read based on thread-based execution. For example, during the execution of instructions in code cache 604, one of those instructions may perform a memory read from a memory address of a hardware device, or from a portion of system memory 608 that is considered non-cacheable. Since there is no cache miss in this case, recording component 106a records the value read with respect to that instruction in trace file(s) 108. As illustrated by the double arrows between operation 704 and operation 705, these operations may occur in any order and may occur multiple times during the tracing of a thread.

[0149] As discussed above, the (one or more) trace files 108 can use various techniques (e.g., using instruction counts, CPU cycles, taken branch counts in conjunction with the program counter, memory accesses, etc.) to record when events occur. In some embodiments, using instruction counts (relative or absolute) to identify when events occur can be advantageous because doing so can remove timing considerations during replay. For example, for the purposes of replay, it can be immaterial how long the underlying hardware took to service a cache miss at the time of recording; one can only care that the recording component 106a recorded that memory address "X" had value "Y", and that the information made its way into the cache 603 in time for instruction "Z" to be dispatched. This can significantly simplify the cost for the replay component 106b to accurately replay the trace because no accurate timing needs to be generated. This method also provides the option of including all cache misses in the (one or more) trace files 108, or only including those cache misses that were actually consumed by the processing unit 602. Thus, for example, speculative reads by the processing unit 602 from the system memory 608 as part of an attempt to anticipate the next instruction to be executed need not be traced. However, as mentioned above, in some other embodiments, the recording component 106a can record timing information in the (one or more) trace files 108. Doing so can enable the replay component 106b to expose that timing during replay when desired.

[0150] Specifically, if the (one or more) trace files 108 map each cache entry to the instruction that brought it into the processor 601, then the (one or more) trace files 108 do not need to capture any information when the cache line is evicted from the cache 603 back to the system memory 608. This is because if the code of the thread again needs that data from the system memory 608, it will be re-imported into the cache 603 by the processor 601 at that time. In such a case, the recording component 106a can record this re-import of the data as a new entry into the (one or more) trace files 108.

[0151] As previously mentioned, in some embodiments, the recording component 106a tracks changes in the control flow of a thread that cannot be determined solely by the code of the thread. For example, when recording user-mode code, there may be additional inputs to the trace because the operating system kernel can interrupt the user-mode code and thus become an external source of data. For example, when certain types of exceptions occur, they are first handled by the kernel and then ultimately control is returned to user mode - this discontinuity is an input into the system. Additionally, when a kernel call is made that does not "return" to an instruction after a system call, this is also a discontinuity into the system. Further, when the kernel changes processor registers before returning control to user-mode code, this is also an input into the system. Thus, the recording component 106a creates records of these discontinuities in the (one or more) trace files 108 such that the replay component 106b reproduces their side effects during replay. Similarly, when recording kernel-mode code (e.g., using a hypervisor), other external inputs (e.g., traps, interrupts, etc.) can occur and the recording component 106a can create records of these discontinuities in the (one or more) trace files.

[0152] In some embodiments, these discontinuities are recorded in the (one or more) trace files 108 as a burst of information. The specific manner in which these discontinuities are recorded can vary as long as the (one or more) trace files 108 contain all of the inputs into the code being recorded to account for the discontinuities. For example, when returning from kernel mode to user mode (e.g., when user-mode code is being recorded), the (one or more) trace files 108 can contain the set of registers (or all registers) that have been changed by the kernel. In another example, the (one or more) trace files include an identification of the address of the continue instruction after a system call. In yet another example, the recording component 106a can flush the processor cache 603 upon returning to user mode, logging all valid cache entries, and / or logging entries only if and when they are used. Similarly, when recording kernel-mode code, the (one or more) trace files can include a record of the continue instruction after a trap / interrupt, any changes made by the trap / interrupt, etc.

[0153] In some embodiments, the recording component 106a and the replay component 106b are configured to track and simulate a processor's translation lookaside buffer ("TLB"), such as Figure 6 The TLB 606 in []. Specifically, it should be noted that the processor cache 603 is sometimes based on the memory physical address of the system memory 608, while the code often refers to the memory virtual address represented by the kernel as an abstraction of the thread / process. The entries in the TLB 606 (e.g., 606a - 606n) store some of the most recent translations between the virtual address and the physical address. Thus, in some embodiments, the recording component 106a records each new entry in the TLB 606 into the (one or more) trace files, which provides all the data required for the replay component 106b to perform the translations during replay. Regarding the data cache 605, the recording component 106a does not need to record any evictions from the TLB 606.

[0154] Tracking TLB 606 entries also allows several other benefits to be provided. For example, the TLB 606 enables the recording component 106a to know which memory pages are uncached or unavailable, so that reads to these pages can be logged (e.g., action 706). Additionally, the TLB 606 enables the recording component 106a to consider cases where two (or more) virtual addresses map to the same physical address. For example, if the cache 603 evicts a physical address entry based on a first virtual address, and then the thread accesses the overlapping physical address via a different virtual address, the recording component 106a uses the TLB 606 to determine the part of the access that it needs to log as a cache miss.

[0155] Some embodiments include a hardware-assisted model that can further reduce the performance impact of cache-based tracing. As previously indicated, recording the trace starts with an empty cache 603 (and thus the cache 603 is flushed). Additionally, if there is a change in the control flow (e.g., due to an interruption), the cache 603 can also be flushed as part of the transition from non-recording (i.e., when in kernel mode) to recording (e.g., when the execution has transitioned back to user mode). However, it should be noted that flushing the processor cache 603 can be computationally expensive. Thus, the embodiments include hardware modifications to the processor 601 that can prevent the need to flush the cache 603 every time the processor 601 transitions from non-recording to recording. These examples refer to set / clear bits. It will be recognized that depending on the implementation, a bit can be considered set to one or zero, and is "cleared" by switching it to the opposite value.

[0156] Some embodiments utilize one or more additional bits that can be used to indicate the state of a cache line to extend each cache line entry (e.g., 604a - 604n of code cache 604 and / or 605a - 605n of data cache 605). For example, one bit can be set (e.g., to one) to indicate that the cache line has been logged to the trace file(s) 108. Then, instead of flushing the cache on a transition, processor 601 only needs to toggle (e.g., to zero) these bits on the cache line. Later, if the code of a thread later consumes a cache line with the bit unset, the entry needs to be stored into the trace as if it were a cache miss. Setting and unsetting the bit also enables recording component 106a to avoid tracing any speculative memory accesses made by processor 601 until it knows that the thread's code has indeed consumed those accesses.

[0157] One or more bits indicating the state of each cache line can also be used in other ways. For example, assume that recording component 106a is recording user - mode code and such code calls into the kernel and then returns to user - mode later. In this case, instead of clearing each cache line bit when processor 601 returns to user - mode, processor 601 can instead clear only the bits on the cache lines modified by the kernel - mode code and leave the bits set on any cache lines that were not modified. This further reduces the amount of entries that recording component 106a needs to add to the trace file(s) 108 when returning from the kernel. However, this technique may not apply to all kernel - to - user - mode transitions. For example, if the kernel simply switches from a thread on another process to the thread being recorded, the bits should be cleared across all cache entries.

[0158] In particular, the foregoing concepts can be used in an environment that uses hyper - threading (i.e., multiple hardware threads are executed on the same processing unit / core). Many processors have a private cache for each core (e.g., each core can have its own level - 1 and level - 2 caches), and also provide a shared cache for all cores (e.g., level - 3 cache). Thus, when using hyper - threading, multiple hardware threads (e.g., "thread 1" and "thread 2") executing on a core share the same private cache (e.g., level - 1, level - 2) for that core, and thus the processor tags each cache line used by these threads with which thread it is associated.

[0159] In some embodiments, one or more bits may be added to each cache line in the shared cache, which indicate whether the cache line has actually been modified by a thread. For example, assume that "Thread 1" is being traced and the execution switches to "Thread 2" on the same core. If "Thread 2" accesses a cache line assigned to "Thread 1" in the private cache of that core, the bit may be switched only if "Thread 2" has modified the cache line. Then, when the execution switches back to "Thread 1" and "Thread 1" accesses the same cache line, the line does not need to be recorded in the trace without "Thread 2" switching the bit, because the content of the cache line has not changed since "Thread 1" last accessed it. The foregoing can also be applied to the shared cache. For example, each cache line in the level 3 cache may include bits for each thread it serves, which indicate whether the thread has "recorded" the current value in the cache line. For example, each thread may set its bit to one when consuming the line and set all other values to zero when writing to the line.

[0160] In some embodiments, the (one or more) trace files 108 may not record entries in the order in which they were executed when they were recorded. For example, assume that entries are recorded based on the number of instructions since the last discontinuity (including kernel calls) in a thread. In this case, the entries between two discontinuities may be reordered without loss of information. In some embodiments, doing so may support faster recording because the processor makes optimizations for memory access. For example, if the processor executes a sequence of instructions (e.g., A, B, and C) that get entries in the (one or more) trace files 108 but do not have dependencies between them, the order in which they are executed is irrelevant during replay. If instruction A accesses data that is not yet in the cache 603, the processor 601 accesses the system memory 608 at the cost of many processor cycles; but instructions B and C access data that is already in the cache 603, and then they can be executed quickly. However, if the instructions need to be recorded in the (one or more) trace files 108 in order, the recording of the execution of instructions B and C must be held (e.g., in the processor's memory resources) until the execution of instruction A has completed. In contrast, if the instructions are permitted to be recorded out of order, instructions B and C may be written to the (one or more) trace files 108 before the completion of instruction A, thus freeing those resources.

[0161] Accordingly, the foregoing embodiments provide new techniques for recording and replaying traces for time-travel debugging that yield performance improvements of several orders of magnitude over previous attempts, support recording of multithreaded programs in which threads freely run simultaneously across multiple processing units, and yield trace files that are reduced in size by several orders of magnitude compared to previously attempted trace files, among other things. Such improvements greatly reduce the amount of computing resources (e.g., memory, processor time, storage space) required by trace and replay software. Accordingly, the embodiments herein can be used in real-world production environments, which greatly enhances the usability and practicality of time-travel debugging.

[0162] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the claims is not necessarily limited to the features described or to the acts described above, or to the order of the acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0163] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is, therefore, indicated by the claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.< / exceptioncontext>

Claims

1. A computer system, comprising: one or more processors; and one or more computer-readable media having computer-executable instructions stored thereon, the computer-executable instructions being executable by the one or more processors to cause the computer system to record a replayable trace of the execution of a multi-threaded process, the computer-executable instructions including instructions executable to cause the computer system to at least perform the following operations: Identify a trace memory model that defines one or more sortable event types that can be sorted across multiple threads of the multi-threaded process, and one or more unsortable event types that cannot be sorted across the multiple threads; Simultaneously execute the multiple threads across one or more processing units of the one or more processors; and During the execution of the multiple threads, independently record a separate replayable trace for each thread, including for each thread: Record the initial state of the thread; Record at least one memory read performed by at least one processor instruction, the at least one processor instruction executed by the thread taking the memory as input; Based on the thread executing an event having a sortable event type, record a first monotonically increasing value into the replayable trace, the monotonically increasing value sorting the event relative to another sortable event, the another sortable event being identified by a second monotonically increasing value recorded in another replayable trace for another thread of the multiple threads; and Record full keyframes and lightweight keyframes, and wherein the frequency of recording the full keyframes or the lightweight keyframes is adjustable.

2. The computer system according to claim 1, wherein recording the initial state for the thread includes: Before executing the thread, record the initial state of one or more processor registers.

3. The computer system according to claim 1, wherein independently recording a separate replayable trace for each thread includes: For at least one thread, record one or more side effects of at least one non-deterministic processor instruction executed by the thread, wherein the side effects include changes made to one or more processor registers by the non-deterministic processor instruction.

4. The computer system according to claim 1, wherein recording at least one memory read performed by at least one processor instruction includes: Based on the at least one memory read not being predicted by a prediction algorithm, record the at least one memory read.

5. The computer system according to claim 1, wherein recording at least one memory read performed by at least one processor instruction includes: Use a shadow copy of the memory to predict the read value, the shadow copy including only a subset of the memory that can be addressed by the process.

6. The computer system according to claim 1, wherein recording at least one memory read performed by at least one processor instruction includes: Record the identity of the memory page associated with the memory read.

7. The computer system according to claim 1, wherein recording an initial state for the thread includes: Record a full keyframe into the replayable trace.

8. The computer system according to claim 1, wherein independently recording a separate replayable trace for each thread comprises: For at least one thread, determine that the memory read does not need to be recorded based on the memory read having been predicted by a prediction algorithm.

9. The computer system according to claim 1, wherein independently recording a separate replayable trace for each thread includes: Record for each thread into one or more reusable circular buffers.

10. The computer system according to claim 1, wherein independently recording a separate replayable trace for each thread includes: Record for each thread one or more interruptions to the execution of the thread.

11. The computer system of claim 1, wherein the trace memory model defines the one or more sortable event types to include at least one of the following: Kernel calls; Full fences; Acquires and releases; or Memory loads.

12. A method implemented at a computer system including one or more processors for recording a replayable trace of the execution of a multi-threaded process, the method comprising: Identifying a trace memory model that defines one or more sortable event types that can be sorted across multiple threads of the multi-threaded process, and one or more unsortable event types that cannot be sorted across the multiple threads; Simultaneously executing the multiple threads across one or more processing units of the one or more processors; And During the execution of the multiple threads, independently recording a separate replayable trace for each thread, including for each thread: Recording an initial state for the thread; Recording at least one memory read performed by at least one processor instruction, where the at least one processor instruction executed by the thread takes the memory as an input; Based on the thread executing an event having a sortable event type, recording a first monotonically increasing value into the replayable trace, the monotonically increasing value sorting the event relative to another sortable event, the another sortable event being identified by a second monotonically increasing value recorded in another replayable trace for another thread among the multiple threads; And Recording full keyframes and lightweight keyframes, and wherein the frequency of recording the full keyframes or the lightweight keyframes is adjustable.

13. The method according to claim 12, wherein recording an initial state for the thread comprises: Before executing the thread, recording an initial state of one or more processor registers.

14. The method according to claim 12, wherein independently recording a separate replayable trace for each thread comprises: For at least one thread, recording one or more side effects of at least one non-deterministic processor instruction executed by the thread, where the side effects include changes made to one or more processor registers by the non-deterministic processor instruction.

15. The method according to claim 12, wherein recording at least one memory read performed by at least one processor instruction comprises: Based on the at least one memory read not being predicted by a prediction algorithm, recording the at least one memory read.

16. The method according to claim 12, wherein recording at least one memory read performed by at least one processor instruction comprises: Using a shadow copy of the memory to predict read values, where the shadow copy includes only a subset of the memory that can be addressed by the process.

17. The method according to claim 12, wherein recording the initial state for the thread comprises: Recording a full keyframe into the replayable trace.

18. The method according to claim 12, wherein independently recording a separate replayable trace for each thread comprises: For at least one thread, determining that the memory read does not need to be recorded based on the memory read having been predicted by a prediction algorithm.

19. The method according to claim 12, wherein independently recording a separate replayable trace for each thread comprises: Recording for each thread into one or more reusable circular buffers.

20. One or more hardware storage devices storing computer-executable instructions that can be executed by one or more processors of a computer system to cause the computer system to record a replayable trace of the execution of a multi-threaded process by causing the computer system to perform the following operations: Identifying a trace memory model that defines one or more sortable event types that can be sorted across multiple threads of the multi-threaded process, and one or more unsortable event types that cannot be sorted across the multiple threads; Simultaneously executing the multiple threads across one or more processing units of the one or more processors; And During the execution of the multiple threads, independently recording a separate replayable trace for each thread, including for each thread: Record the initial state for the thread; Record at least one memory read performed by at least one processor instruction, where the at least one processor instruction executed by the thread takes the memory as an input; Based on the thread executing an event having a sortable event type, record a first monotonically increasing value into the replay trace, where the monotonically increasing value sorts the event relative to another sortable event, the another sortable event being identified by a second monotonically increasing value, the second monotonically increasing value being recorded in another replay trace for another thread among the plurality of threads; And Record full key frames and lightweight key frames, and where the frequency of recording the full key frames or the lightweight key frames is adjustable.

Citation Information

Patent Citations

  • System and method for debugging of computer programs

    EP2600252A1

  • Method and framework for tracking / logging completion of requests in a computer system

    US20050021708A1