Data Race Analysis Based on Changing Data Loaded Inside Functions During Time Travel Debugging

By replacing the load value during function playback, simulating or correcting data competition, the problem of difficult data competition in multi-threaded applications is solved, and the reliability and debugging efficiency of the code are improved.

CN114245892BActive Publication Date: 2025-07-29MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080056933.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-14
Filing Date
2020-06-11
Publication Date
2025-07-29
Estimated Expiration
2040-06-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and fix data competition problems in multithreaded applications, especially when performing timing characteristics changes, which are difficult to reproduce and confirm.

Method used

Use historical debugging technology to replace the load value in the function during function playback, and observe its impact on function output by simulating or correcting data competition, and determine the occurrence and potential impact of memory competition.

Benefits of technology

Simplifies the process of identifying the root causes of data competition, reduces development, troubleshooting and debugging time, and improves code reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114245892B_ABST
    Figure CN114245892B_ABST
Patent Text Reader

Abstract

Determine whether the loading modification within the function affects the output of the traced function. The function is identified within the traced portion of the previous execution of the entity. The function includes a sequence of executable instructions and produces one or more outputs. The (multiple) traced output data values produced by the traced instance of the function are determined, and the executable instructions loaded from memory during the execution within the sequence of executable instructions are identified. The execution of the function is emulated while replacing the traced memory values loaded by the executable instructions during the traced instance of the function with different memory values, and while producing (multiple) emulated output data values. Based on a difference existing between the (multiple) traced output data values and the (multiple) emulated output data values, a notification is generated at the user interface or to a software component.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Tracking and correcting undesired software behavior is a core activity in software development. Undesired software behavior can include many things, such as execution crashes, runtime exceptions, slow execution performance, incorrect data results, data corruption, etc. Undesired software behavior can be triggered by a variety of factors, such as data input, user input, data races (e.g., when accessing shared resources), etc. Given the diversity of triggers, undesired software behavior can be rare and seemingly random, and extremely difficult to reproduce. Therefore, it can be very time-consuming and difficult for developers to identify a given undesired software behavior. Once an undesired software behavior has been identified, determining its root cause(s) again can be both time-consuming and difficult.

[0002] When developing multithreaded applications, identifying and fixing data races can be particularly time-consuming and difficult. A data race occurs when multiple threads access the same memory location and at least one of the threads writes to that memory location without carefully controlling the order of execution of these threads. Thus, for example, a first thread may write to a memory location when a second thread does not expect it, causing the second thread to read an invalid or unexpected value from the memory location. For example, data races can occur due to careless thread synchronization when deliberately using shared memory (e.g., global variables, heap allocations, etc.), or due to threads inadvertently writing to shared memory (e.g., using rogue pointers, using incorrect offsets with otherwise valid pointers, etc.). Data races can be difficult to identify because they may not be reliably reproducible - their occurrence (or lack thereof) can depend on the timing of how the executions of multiple threads interleave - which can vary over time depending on the workload of each thread, the timing characteristics of accessing memory and other input / output devices, the overall workload of the processor, specific user input, etc.

[0003] Developers have conventionally used a variety of methods to identify undesired software behavior and then identify the location(s) in the application code that cause the undesired software behavior. For example, developers may test different parts of the application code for different inputs (e.g., unit tests). As another example, developers may infer the code execution of the application in a debugger (e.g., by setting breakpoints / watchpoints, stepping through code lines as the code executes, etc.). As another example, developers may observe the code execution behavior in a profiler (e.g., timing, coverage). As another example, developers may insert diagnostic code (e.g., trace statements) into the application code.

[0004] While conventional diagnostic tools (e.g., debuggers, profilers, etc.) operate on code executing forward "in the live", emerging forms of diagnostic tools enable "historical" debugging (also known as "time travel" or "reverse" debugging), in which the execution of at least a portion of one or more program threads is recorded into one or more trace files (i.e., the recorded execution). Using some tracing techniques, the recorded execution can include "bit-accurate" historical trace data, which enables the recorded portions of the one or more traced threads to be virtually "replayed" down to the granularity of individual instructions (e.g., machine code instructions, intermediate language code instructions, etc.). Thus, using the "bit-accurate" trace data, diagnostic tools can enable developers to reason about the previously recorded execution of the subject code rather than executing the code forward "in the live". For example, a historical debugger can enable forward and reverse breakpoints / watchpoints, can enable the code to be single-stepped forward and backward, etc. On the other hand, a historical profiler can be able to derive code execution behavior (e.g., timing, coverage) from previously executed code. Summary of the Invention

[0005] At least some embodiments described herein utilize historical debugging techniques to replace loaded values within a function during function replay in order to observe the effect of the replacement on the output(s) of the function. Thus, according to embodiments herein, when simulating the execution of a given function, the emulator is able to modify one or more memory values read by the function and then compare the simulated behavior of the function to the traced behavior. In some embodiments, the debugger uses the comparison to determine whether a memory race has occurred (or is likely to occur) during the traced function and / or to identify the potential effects of the memory race. In some embodiments, the debugger can perform this analysis regardless of whether a data race actually occurred during the trace.

[0006] For example, modifying the value read by a given load can be operative to simulate the occurrence of a data race. Thus, the debugger can observe the effect that the simulated data race might have on the behavior of the function (e.g., based on the value(s) of the output(s) of the function after the simulated data race), even if no data race actually occurred during the trace. Thus, when a race can be suspected but was not captured during the trace, the debugger can implement a simulation of the data race in the function. Alternatively, modifying the value read by a given load can be operative to simulate the correction of a data race. Thus, when a data race was captured during the trace, the debugger can observe the effect that the simulated correction of the data race might have on the behavior of the function (e.g., based on the value(s) of the output(s) of the function after the simulated correction of the data race).

[0007] In some embodiments, a method, system, and computer program product use recorded execution to determine whether a load modification within a function affects one or more outputs of a traced function. In these embodiments, a computer system accesses recorded execution that includes trace data of a previous execution of at least a portion of the executable code of an executable entity. The trace data enables replaying the previous execution of the portion of the executable entity. The computer system identifies a function within the traced portion of the executable code of the executable entity. The function includes a sequence of executable instructions that consume zero or more inputs and produce one or more outputs. Based on the trace data, the computer system determines one or more traced output data values that were produced by a traced instance of the function during the previous execution. The computer system identifies at least one executable instruction that performs a load from memory within the sequence of executable instructions of the function. The computer system emulates the execution of the function according to the trace data. The emulation includes replacing the traced memory values loaded by the at least one executable instruction during the traced instance of the function with different memory values and producing one or more emulated output data values for the one or more outputs. The computer system determines whether there is a difference between the one or more traced output data values and the one or more emulated output data values. Based on there being a difference between the one or more traced output data values and the one or more emulated output data values, the computer system generates a notification at a user interface or to a software component.

[0008] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to assist in determining the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To describe the manner in which the above and other advantages and features of the present invention can be obtained, a more particular description of the invention briefly described above will be presented by reference to specific embodiments thereof that are illustrated in the accompanying drawings. It is to be understood that these drawings depict only typical embodiments of the invention and are not to be considered limiting of its scope, for the invention will be described and explained with additional specificity and detail through the use of the accompanying drawings, in which:

[0010] Figure 1A An example computing environment that facilitates using recorded execution to determine whether a load modification within a function affects the output of a traced function is illustrated;

[0011] Figure 1B An example debug component is illustrated;

[0012] Figure 2 An example computing environment is illustrated in whichFigure 1A A computer system is connected to one or more other computer systems via one or more networks;

[0013] Figure 3 Illustrates an example of the recorded execution of an executable entity;

[0014] Figure 4 Illustrates an example of a function in an executable entity, where the function is identified based on its input and output;

[0015] Figures 5A to 5C Illustrates an example of a potential data race situation; and

[0016] Figure 6 Illustrates a flowchart of an example method for using recorded execution to determine whether a load modification inside a function affects the output of the traced function. Detailed Description

[0017] At least some embodiments described herein utilize historical debugging techniques to replace loaded values within a function during function replay in order to observe the effect of the replacement on the output(s) of the function. Thus, according to embodiments herein, when simulating the execution of a given function, the emulator is able to modify one or more memory values read by the function and then compare the simulated behavior of the function to the traced behavior. In some embodiments, the debugger uses this comparison to determine whether a memory race has occurred (or may occur) during the tracing of the function, and / or to identify the potential impact of the memory race. In some embodiments, the debugger can perform this analysis regardless of whether a data race actually occurs during the tracing.

[0018] For example, modifying the value read by a given load can be used to simulate the occurrence of a data race. Thus, the debugger can observe the effect that the simulated data race may have on the behavior of the function (e.g., based on the value(s) of the output(s) of the function after the simulated data race), even if no data race actually occurred during the tracing. Thus, when a race can be suspected but not captured during the tracing, the debugger can implement data race simulation in the function. Alternatively, modifying the value read by a given load can be used to simulate the correction of a data race. Thus, when a data race is captured during the tracing, the debugger can observe the effect that the simulated correction of the data race may have on the behavior of the function (e.g., based on the value(s) of the output(s) of the function after the simulated correction of the data race).

[0019] As will be appreciated by those skilled in the art, replacing loaded values within a function during function replay and observing the effect of the replacement on the output(s) of the function can provide many technical benefits. For example, it can be very difficult to reproduce a data race during testing. Additionally, even if a data race occurs in production, the data race may not be reproducible during tracing due to the changes in the execution timing characteristics introduced by the tracing. Thus, even if a program error that can be suspected of being a data race occurs, it may be difficult, if not impossible, to capture the trace of the data race. Embodiments herein provide tools that enable a suspected data race to be simulated. Based on the simulation of the data race, the suspected data race can be confirmed, thus greatly simplifying the process of identifying the root cause of the data race. Additionally, even if a data race has occurred during tracing, it may not be clear whether a data race has actually occurred. By simulating the correction of the data race and observing the resulting behavior, data race confirmation can be greatly simplified. Additionally, even if a data race has occurred during tracing, it may not be clear what effect correcting the data race may have. By simulating the correction of the data race, these effects can be observed via emulation - even before a code fix is generated. The net effect of these technical benefits is to produce less error-prone code, reducing the time spent on development, troubleshooting, and debugging.

[0020] As indicated, embodiments herein operate on the recorded execution of an executable entity. In this specification and the appended claims, "recorded execution" can refer to any data that stores a record of a previous execution of (a plurality of) code instructions, or that can be used to at least partially reconstruct the previous execution of (a plurality of) previously executed code instructions. Generally, these code instructions are part of an executable entity and are executed on (a plurality of) physical or virtual processors as threads and / or processes (e.g., as machine code instructions), or in a managed runtime (e.g., as intermediate language code instructions).

[0021] The recorded execution used by embodiments herein can be generated by various historical debugging techniques. Generally, historical debugging techniques record or reconstruct the execution state of an entity at various times so that the execution of the entity can be at least partially emulated from that execution state at a later time. The fidelity of this virtual execution varies depending on the available recorded execution state.

[0022] For example, one class of historical debugging techniques (referred to herein as time travel debugging) continuously records a bit-exact trace of an entity's execution. This bit-exact trace can then later be used to faithfully replay the entity's previous execution down to the fidelity of individual code instructions. For example, the bit-exact trace can record information sufficient to reproduce the initial processor state at at least one point in the thread's previous execution (e.g., by recording a snapshot of the processor registers) and the data values read when the thread's instructions execute after that point in time (e.g., memory reads). This bit-exact non-day trace can then be used to replay the execution of the thread's code instructions (starting with the initial processor state) by supplying the recorded reads to the instructions.

[0023] Another class of historical debugging techniques (referred to herein as branch trace debugging) relies on working backward from a dump or snapshot (e.g., a thread's crash dump) that includes a processor branch trace (i.e., a record of whether branches were taken) to reconstruct at least a portion of an entity's execution state. These techniques start with the values from the dump or snapshot (e.g., memory and registers) and use the branch trace to at least partially determine the code execution flow, iteratively replay the entity's code instructions, and work backward and forward to reconstruct the intermediate data values (e.g., registers and memory) used by the code until those values reach a stable state. These techniques may be limited in how far back they can reconstruct data values and how many data values can be reconstructed. Nevertheless, the reconstructed historical execution data can be used for historical debugging.

[0024] Yet another class of historical debugging techniques (referred to herein as replay and snapshot debugging) periodically records a complete snapshot of an entity's memory space and processor registers while the entity is executing. If the entity depends on data from sources other than the entity's own memory or from non-deterministic sources, these techniques can also record such data along with the snapshot. These techniques then use the data in the snapshot to replay the execution of the entity's code between snapshots.

[0025] Figure 1A Illustrated is an example computing environment 100a that facilitates using recorded execution to determine whether a load modification within a function affects the output of a traced function (which may be useful for data race analysis, for example). As depicted, computing environment 100a can include or utilize a special-purpose or general-purpose computer system 101 that includes computer hardware such as, for example, one or more processors 102, system memory 103, persistent storage device 104, and / or (multiple) network devices 105 that are communicatively coupled using one or more communication buses 106.

[0026] Embodiments within the scope of the present invention may include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media accessible by a general or special-purpose computer system. A computer-readable medium storing computer-executable instructions and / or data structures is a computer storage medium. A computer-readable medium carrying computer-executable instructions and / or data structures is a transmission medium. Thus, by way of example and not limitation, embodiments of the present invention may include at least two distinct types of computer-readable media: computer storage media and transmission media.

[0027] Computer storage media are physical storage media (such as system memory 103 and / or persistent storage device 104) that store computer-executable instructions and / or data structures. Physical storage media include computer hardware, such as RAM, ROM, EEPROM, solid state drives (“SSD”), flash memory, phase change memory (“PCM”), optical disk storage devices, magnetic disk storage devices, or any other hardware storage device(s) that can be used to store program code in the form of computer-executable instructions or data structures that can be accessed and executed by a general or special-purpose computer system to implement the functionality of the present invention as disclosed.

[0028] Transmission media can include a network and / or a data link that can be used to carry program code in the form of computer-executable instructions or data structures and can be accessed by a general or special-purpose computer system. A “network” is defined as one or more data links capable of transporting electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer system via a network or another communication connection (wired, wireless, or a combination of wired or wireless), the computer system can consider the connection to be a transmission medium. Combinations of the above should also be included within the scope of computer-readable media.

[0029] Further, when program code in the form of computer-executable instructions or data structures reaches various computer system components, it can be automatically transferred from the transmission medium to the computer storage medium (and vice versa). For example, computer-executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (such as (one or more) network devices 105) and then ultimately transferred to the computer system RAM (such as system memory 103) and / or a less volatile storage medium (such as persistent storage device 104) at the computer system. Thus, it should be understood that computer-readable media can be included in computer system components that also (or even primarily) use transmission media.

[0030] For example, computer-executable instructions include instructions and data that, when executed at one or more processors, cause a general-purpose computer system, a special-purpose computer system, or a special-purpose processing device to perform a particular function or group of functions. The computer-executable instructions can be, for example, machine code instructions (such as binary), intermediate format instructions (such as assembly language), or even source code.

[0031] Those skilled in the art will appreciate that the present invention can be practiced in a network computing environment using many types of computer system configurations, including personal computers, desktop computers, laptop computers, messaging processors, handheld devices, multiprocessor systems, multiprocessor- or programmable-based consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, and the like. The present invention can also be practiced in a distributed system environment where both local and remote computer systems that are linked (by hardwired data links, wireless data links, or a combination of hardwired and wireless data links) through a network perform tasks. Thus, in a distributed system environment, the computer system can include multiple constituent computer systems. In a distributed system environment, program modules can be located in both local and remote memory storage devices.

[0032] Those skilled in the art will also appreciate that the present invention can be practiced in a cloud computing environment. The cloud computing environment can be distributed, although this is not required. When distributed, the cloud computing environment can be distributed internationally within an organization and / or have components that are processed across multiple organizations. In this specification and the following claims, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources (such as networks, servers, storage devices, applications, and services). The definition of "cloud computing" is not limited to any of the many other advantages that can be obtained from such a model when appropriately deployed.

[0033] The cloud computing model can consist of various characteristics, such as on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so on. The cloud computing model can also appear in the form of various service models, such as, for example, software as a service ("SaaS"), platform as a service ("PaaS"), and infrastructure as a service ("IaaS"). The cloud computing model can also be deployed using different deployment models, such as private cloud, community cloud, public cloud, hybrid cloud, and so on.

[0034] Some embodiments, such as cloud computing environments, may include a system that includes one or more hosts, each capable of running one or more virtual machines. During operation, the virtual machines emulate an operable computing system, thereby supporting an operating system and possibly one or more other applications. In some embodiments, each host includes a hypervisor that emulates the virtual resources of the virtual machines using the physical resources abstracted from the view of the virtual machines. The hypervisor also provides appropriate isolation between the virtual machines. Thus, from the perspective of any given virtual machine, the hypervisor provides the illusion that the virtual machine is interfacing with the physical resources, even though the virtual machine is only interfacing with the appearance of the physical resources (e.g., virtual resources). Examples of physical resources include processing power, memory, disk space, network bandwidth, media drives, and the like.

[0035] As Figure 1A shown, each processor 102 may include, among other things, one or more processing units 107 (e.g., processor cores) and one or more caches 108. Each processing unit 107 loads and executes machine code instructions via the cache 108. During the execution of these machine code instructions at one or more execution units 107b, the instructions may use the internal processor registers 107a as temporary storage locations and may read from and write to various locations in the system memory 103 via the cache 108. Generally, the cache 108 temporarily caches portions of the system memory 103; for example, the cache 108 may include a "code" portion that caches portions of the system memory 103 storing application code and a "data" portion that caches portions of the system memory 103 storing application runtime data. If a processing unit 107 needs data (e.g., code or application runtime data) that is not yet stored in the cache 108, the processing unit 107 may initiate a "cache miss", causing the required data to be fetched from the system memory 103 - possibly evicting some other data from the cache 108 back to the system memory 103.

[0036] As illustrated, the persistent storage device 104 may store computer-executable instructions and / or data structures representing executable software components; correspondingly, during the execution of the software at the (one or more) processors 102, one or more portions of these computer-executable instructions and / or data structures may be loaded into the system memory 103. For example, the persistent storage device 104 is shown storing computer-executable instructions and / or data structures corresponding to a debugging component 109, a tracker component 110, an emulation component 111, and one or more applications 112. The persistent storage device 104 may also store data, such as one or more recorded executions 113 (e.g., generated using one or more of the historical debugging techniques described above).

[0037] Typically, the debug component 109 utilizes the emulation component 111 to emulate the execution of code of an executable entity (such as the application(s) 112) based on execution state data obtained from one or more of the recorded execution(s) 113. Thus, Figure 1A is shown with the debug component 109 and the emulation component 111 loaded into the system memory 103 (i.e., debug component 109' and emulation component 111'), and the application(s) 112 being emulated within the emulation component 111' (i.e., application(s) 112').

[0038] Typically, the tracer component 110 records or "traces" the execution of one or more of the application(s) 112 into the recorded execution(s) 113 (e.g., using one or more of the historical debugging techniques described above). The tracer component 110 can record the execution of the application(s) 112 whether the execution is a "live" execution directly on the processor(s) 102, or the execution is a "live" execution on the processor(s) 102 via a managed runtime, and / or the execution is an emulated execution via the emulation component 111. Thus, Figure 1A is also shown with the tracer component 110 loaded into the system memory 103 (i.e., tracer component 110'). The arrow between the tracer component 110' and the recorded execution(s) 113' indicates that the tracer component 110' can record trace data into the recorded execution(s) 113' which can then be persisted as the recorded execution(s) 113 to the persistent storage device 104.

[0039] The computer system 101 can additionally or alternatively receive one or more of the recorded execution(s) 113 from another computer system (e.g., using the network device(s) 105). For example, Figure 2 illustrates an example computing environment 200 where Figure 1A the computer system 101 is connected to one or more other computer systems 202 (i.e., computer systems 202a through 202n) via one or more networks 201. As shown in example 200, each computer system 202 includes a tracer component 110 and the application(s) 112. Thus, the computer system 101 can receive one or more of the recorded execution(s) 113 of one or more of the application(s) 112 at these computer systems 202 via the network(s) 201.

[0040] Note that, while the debugging component 109, the tracer component 110, and / or the emulation component 111 may each be separate components or applications, they may alternatively be integrated into the same application (such as a debugging suite), or may be integrated into another software component - such as an operating system component, a hypervisor, a cloud infrastructure, etc. Thus, those skilled in the art will also appreciate that the present invention may be practiced in a cloud computing environment of which the computer system 101 is a part.

[0041] As previously mentioned, the debugging component 109 utilizes the emulation component 111 to emulate the execution of the code of one or more applications 112 using execution state data from one or more of the recorded executions 113. According to embodiments herein, when emulating the execution of a given application 112, the debugging component 109 is capable of modifying one or more memory values read by the application 112 and then comparing the emulated behavior to the traced behavior. Based on this comparison, the debugging component 109 can determine whether a memory contention has occurred (or is likely to occur) during the recorded execution of the application 112 and / or identify the potential impact of the memory contention, regardless of whether the contention actually occurred during tracing. In some embodiments, modifying the memory values can operate to simulate the occurrence of a memory contention (e.g., when the contention is suspected but not captured during tracing), or to simulate the effect of correcting a memory contention (e.g., when the contention is captured during tracing). As will be explained, embodiments herein can be used to detect a memory contention (or a possible memory contention), even when the execution of the entity causing the memory contention is not traced to one or more of the recorded executions 113. In some embodiments, the debugging component 109 may also utilize the tracer component 110 to record the emulated execution into the recorded executions 113 for further historical debugging analysis (e.g., by adding additional trace data to the existing recorded executions 113, and / or by creating new recorded executions 113).

[0042] To demonstrate how the debugging component 109 can accomplish the foregoing, Figure 1B illustrates an example 100b providing additional details of the debugging component 109 that Figure 1A provides. Figure 1BThe debug component 109 depicted includes various components (such as data access 114, trace / code analysis 115, simulation 116, simulation analysis 117, output 118, etc.) that represent various functions that the debug component 109 can be implemented according to the various embodiments described herein. It is to be understood that the depicted components - including their identities, sub-components, and arrangements - are presented only to assist in describing the various embodiments of the debug component 109 described herein, and these components are not limited to the various embodiments of how software and / or hardware can implement the debug component 109 described herein or its specific functionality.

[0043] As shown, the data access component 114 includes a trace access component 114a and a code access component 114b. The trace access component 114a accesses one or more of the recorded executions 113 in the (multiple) recorded executions, such as the recorded executions 113 of a previous execution of the application 112. Figure 3 An example of a recorded execution 300 of an executable entity (such as the application 112) is illustrated, which can be accessed by the trace access component 114a, where the recorded execution 300 may have been generated using time travel debugging techniques.

[0044] In Figure 3 the example of, the recorded execution 300 includes a plurality of data streams 301 (i.e., data streams 301a to 301n). In some embodiments, each data stream 301 records the execution of a different thread from the code execution of the application 112. For example, the data stream 301a may record the execution of the first thread of the application 112, while the data stream 301n records the nth thread of the application 112. As shown, the data stream 301a includes a plurality of data packets 302. Since the specific data in each data packet 302 of the recorded log may be different, they are shown as having different sizes. Generally, when using time travel debugging techniques, each data packet 302 records at least the input (such as register values, memory values, etc.) to one or more executable instructions executed as part of this first thread of the application 112. As shown, the data stream 301a may also include one or more key frames 303 (such as key frames 303a to 303b), each of which records sufficient information (such as a snapshot of register and / or memory values) such that the previous execution of the thread can be replayed by the simulation component 116 starting from a point forward of the key frame.

[0045] In some embodiments, the recorded execution 113 may also include the actual code executed as part of the application 112. Thus, in Figure 3In this, each data packet 302 is shown as including a non-shaded data input portion 304 and a shaded code portion 305. In some embodiments, the code portion 305 (if any) of each data packet 302 may include executable instructions executed based on the corresponding data input. However, in some other embodiments, the recorded execution 113 may omit the actual code being executed and instead rely on having separate access to the code of the application 112 (such as from the persistent storage device 104). In these other embodiments, each data packet may specify, for example, the address or offset of the appropriate executable instruction(s) in the application binary image. Although not shown, the recorded execution 300 may also include a data stream 301 that stores one or more of the outputs of the code execution.

[0046] If there are multiple data streams 301, each recording the execution of a different thread, these data streams may include ordering events. Each ordering event records the occurrence of an event that can be ordered across threads. For example, an ordering event may correspond to an interaction between threads, such as an access to memory shared by the threads. Thus, for example, if a first thread traced to a first data stream (such as 301a) writes to a synchronization variable, a first ordering event may be recorded in that data stream (such as 301a). Later, if a second thread traced to a second data stream (such as 301b) reads from that synchronization variable, a second ordering event may be recorded in that data stream (such as 301b). These ordering events may be inherently ordered. For example, each ordering event may be associated with a monotonically increasing value, where the monotonically increasing value defines the total order between the ordering events. For example, the first ordering event recorded in the first data stream may be given the value 1, the second ordering event recorded in the second data stream may be given the value 2, and so on.

[0047] Returning to Figure 1B , the code access component 114b may obtain the code of the application 112. If the recorded execution(s) 114 obtained by the trace access component 114a include the traced code (such as the code portion 305), the code access component 114b may extract the code from the recorded execution 113. Alternatively, the code access component 114b may obtain the code of the application 112 from the persistent storage device 104 (such as from the application binary image).

[0048] The trace / code analysis component 115 can perform one or more types of analysis on the recorded execution 113 and / or the application 112 accessed by the data access component 114, which supports modifying (multiple) memory values previously read by the application 112 during simulation of the recorded execution of the application 112. This data modification can be used, for example, as part of data race analysis. As an aid in describing possible types of analysis, the trace / code analysis component 115 is shown as potentially including a function identification component 115a, an input / output identification component 115b, and a load identification component 115c.

[0049] In some embodiments, the function identification component 115a identifies discrete "functions" in the code of the subject application 112. In some embodiments, the function identification component 115a can identify the functions based on identifying the inputs and outputs from those functions. As used herein, a "function" is defined as a collection of one or more segments of executable code, each segment including a sequence of one or more executable instructions having zero or more "inputs" and one or more "outputs". For example, Figure 4 An example 400 of a function in an executable entity is illustrated, where the function is identified based on its inputs and outputs. Specifically, Figure 4 A representation 401 of the code of the application 112 is shown. Figure 4 Also shown is that different segments of the execution are identified as different functions (i.e., function 402, including functions 402-1 to 402-9) in the representation 401. Each of these functions 402 can have a different set of zero or more inputs and a different set of one or more outputs. For example, in Figure 4 each function 402 has a corresponding set of (multiple) inputs 403 and a corresponding set of (multiple) outputs 404. For example, function 402-1 has a set of (multiple) inputs 403-1 and a set of (multiple) outputs 404-1, function 402-2 has a set of (multiple) inputs 403-2 (which may correspond to the set of outputs 404-1) and a set of outputs 404-2, and so on. Thus, in conjunction with the function identification component 115a identifying the functions, the input / output identification component 115b can identify the corresponding inputs and outputs.

[0050] Although the scope of the identified functions may vary, in some embodiments, the function identification component 115a can identify functions based on identifying basic blocks in the application 112. As will be understood by one of ordinary skill in the relevant art and as used herein, a "basic block" is a sequence of instructions that serves as an execution unit; that is, the sequence has a single input point and a single output point, and all or none of the instructions in the basic block execute or do not execute (except for exceptions). Thus, in some embodiments, a "function" can include a single basic block, although a function can alternatively include multiple basic blocks.

[0051] As used herein, an "input" is defined as any data location from which a function (as defined above) reads and which the function itself has not written to prior to the read. These data locations can include, for example, registers that exist at the time the function is entered and / or any memory locations from which the function reads but which it has not allocated itself. Edge cases can occur if the function allocates memory and then reads from that memory prior to initialization. In these instances, embodiments can treat reads of uninitialized memory as either inputs or bugs. As used herein, an "output" is defined as any data location (e.g., register and / or memory location) to which a function writes and which it does not deallocate later. For example, a stack allocation at the entry of a function, followed by a write to the allocated region, followed by a stack deallocation at the exit of the function, is not considered an output of the function. Additionally, if a function is bounded by an application binary interface (ABI) boundary, any volatile registers (i.e., registers not used to pass the return value) at the exit of the function are implicitly "deallocated" (i.e., they are discarded by the ABI) - and thus are not outputs of the function.

[0052] In some embodiments, the function identification component 115a may rely on the operating system and the (multiple) applications 112 being compiled for a processor instruction set architecture (ISA) with a known ABI - which reduces the need to separately track registers. Thus, for example, instead of separately tracking registers, the function identification component 115a can use the ABI for which the (multiple) applications 112 are compiled to determine which (multiple) registers the (multiple) applications 112 use to pass parameters to a function, and / or which (multiple) registers the (multiple) applications 112 use for return values. In some embodiments, debug symbols can be used to supplement or replace ABI information. Notably, even if a calling function ignores the return value of a called function, the ABI and / or symbols can still be used to determine whether the contents of the register used to store the return value of the called function have changed.

[0053] In some embodiments, the function identification component 115a may define and map functions that include instruction sequences having one or more gaps within their execution. For example, a function may include an instruction sequence that makes a kernel call (which may not be logged) in the middle of execution. For illustration, function 402-1 may take a file handle and a character as inputs and include instructions that compare each byte of the file with the input character to find the occurrence of the character in the file. Since they depend on file data, these instructions may make one or more kernel calls (not shown) to read the file (e.g., using the handle as an argument to the kernel call). To identify functions with gaps, the function identification component 115a may need to ensure that these gaps are correctly ordered within each function with respect to comparison operations so that the file data is processed in the same order in each function. Note that, in some embodiments, any register values changed by a kernel call are typically tracked in the (multiple) recorded executions 113. Nevertheless, the function identification component 115a may additionally or alternatively use the ABI and / or debug symbols to track which register values are preserved across kernel calls. For example, the stack pointer (i.e., ESP on x86 or R13 on ARM) may be preserved across kernel calls.

[0054] In some embodiments, inputs and outputs are composable. For example, if a single function in application 112 is inclusively defined as the entirety of the code in sections A, B, and C, the input set for that function may be defined as the input set that includes the combination of each input in the inputs of sections A, B, and C, and its output set may be defined as the output set that includes the combination of each output in the outputs of sections A, B, and C. It is to be understood that when an input (or output) of section B is allocated (or deallocated) by section A, or if it is allocated by section B and deallocated by section A, that input (or output) of function B may be omitted from the input set (or output set). It is also to be understood that any input (or output) of a section that is called within a broader function (i.e., that includes that section) and is not an input (or output) of the broader function may be omitted from the input set (or output set) of the broader function or may otherwise be tracked as internal to the broader function.

[0055] Due to function inlining, complex situations can also occur, especially when the sub-function will not be analyzed by the debugging component 109 (e.g., because it comes from a third-party library). For example, assume that the first section (A1) of function A is executed before calling sub-function B, and then the second section (A2) of function A is executed after function B returns. Here, sections A1 and A2 themselves may be regarded as independent functions, which have their own sets of inputs and outputs. If function B takes any output of A1 as an input, then these outputs need to be generated before calling function B; similarly, if function A2 takes any output of function B as an input, then these outputs need to appear after the call to function B.

[0056] In some embodiments, the input / output identification component 115b identifies the actual values of one or more inputs and / or outputs among the inputs and / or outputs of the functions that occur during the tracing. This can include, for example, identifying the values of one or more inputs and / or one or more outputs that are traced to the recorded execution 113 and / or causing one or more functions in the function to be replayed by the emulation component 116 in order to reproduce the values of one or more inputs and / or one or more outputs that occurred during the tracing (e.g., to obtain values that may not have been traced to the recorded execution 113).

[0057] For a given function, the load identification component 115c identifies one or more loads that are executed within the function. The identified loads can load values from a memory location, from a register location, or from some other memory. In some embodiments, the identified load can be addressed to one of the inputs in the function input; thus, it would be a load from the input location before the function has executed a store to the input location. In this case, the load is expected to obtain data external to the function. In other cases, the load can be to data within the function. For example, the load can be addressed to a location that previously corresponded to an input, but after the function has executed one or more stores to that location (therefore, the location is no longer an input at the time of the load). In another example, the load can be addressed to a location that does not correspond to any input, but after the function has executed one or more stores to that location (e.g., a variable / data structure allocated by the function).

[0058] Although the load identification component 115c can identify all loads within a given function, the load identification component 115c can potentially identify only those loads for which the (multiple) values of the (multiple) loads can be traced to the (multiple) values of the function output. Thus, for example, the load identification component 115c can perform a "reverse taint" analysis of the function starting from the output of the function to identify data continuity between one or more loads and the output. For example, starting from a particular output, the load identification component 115c can analyze the sequence of executable instructions of the function to identify the instruction that executes to place the output value of the function into the storage of the output. The load identification component 115c can then identify the executable instructions corresponding to one or more first loads that are used to produce the output value written by the storage. Each of the (multiple) first loads among the (multiple) these first loads can then be identified by the load identification component 115c as a load that contributes to the output. The process can then be repeated for each memory location corresponding to each first load to identify the executable instructions corresponding to a (multiple) second load that contributes to the (multiple) values read by the (multiple) first loads. The process can be repeated any number of times to identify a set of one or more loads that contribute to the value of a particular output.

[0059] In some embodiments, the load identification component 115c operates to identify loads that are likely to read data affected by a data race and omit loads that are unlikely to read data affected by a data race. Thus, the load identification component 115c can use one or more heuristics to classify loads based on the relative likelihood that the target memory of the load can be used by multiple threads. Such heuristics can be aided by debug symbols, if available. For example, if a load within a given function is to a memory location corresponding to the function stack or to a constant, the data read by the load is unlikely to be affected by a data race. Thus, the load identification component 115c can omit that load from the identified loads. In contrast, if a load within a given function is to a memory location corresponding to heap memory or to a global variable, the load may be affected by a data race. Thus, the load identification component 115c can include that load in the identified loads. In some embodiments, stack memory can be traced even more finely. For example, an embodiment can determine whether a pointer to a stack address of a given function has been occupied by another function. If so, that address is likely to be more affected by a data race than another stack address that has not been occupied by another function. Additionally, a function with no stack address occupied by another function is less likely to suffer from a data race than a function with one or more stack addresses occupied by another function. As mentioned, the recorded execution 113 can include ordering events. In some embodiments, the load identification component 115c can consider these ordering events. For example, if a load corresponds to an ordering event, it may be part of a cross-thread synchronization event. Thus, that load is less likely to be part of a data race, and the load identification component 115c may omit that load from the identified loads or at least mark it as less likely to be part of a data race compared to other loads. In some embodiments, the load identification component 115c can operate at least in part based on user input. For example, the load identification component 115c can receive user input specifying a particular load of interest, can receive user input specifying a variable or data structure of interest, etc.

[0060] For each load identified, the load value identification component 115d can identify one or more replacement values for the load. Each replacement value is a value that could be artificially used as the value obtained by the load during simulation of the load by the simulation component 116. The load value identification component 115d can operate separately from and / or in conjunction with the simulation performed by the simulation component 116. Thus, for example, the load value identification component 115d can identify the (multiple) replacement values for one or more loads in a given function before and / or during function simulation.

[0061] The load value identification component 115d can select from a variety of methods for determining what value to substitute for a particular load. For example, the load value identification component 115d can prompt the user for input to obtain a value from the user, or can select a random value. The load value identification component 115d can alternatively perform a function analysis to identify an appropriate substitution value. For example, if trace data for multiple previously executed instances of a function is available, the load value identification component 115d can analyze these multiple instances to identify the values read by each instance of the function for a given load. Then, for a given instance of the function and that load, the load value identification component 115d can select a substitution value that is typical for the load (e.g., if attempting to simulate the correction of a data race), or select a substitution value that is atypical for the load (e.g., if attempting to simulate the introduction of a data race).

[0062] In another example, the load value identification component 115d can perform an analysis of the data flow within a function in conjunction with the trace data to determine what value to substitute. To illustrate this concept, Figures 5A to 5C Examples 500a through 500c of potential data race situations are illustrated. As mentioned in the description of the load identification component 115c, the identified load can be addressed to, for example, a memory location corresponding to a function input, to a memory location that was previously an input, or to a memory location not corresponding to an input.

[0063] Example 500a illustrates a load from a memory location corresponding to a function input. Specifically, rectangle 501a represents the execution of the function over time - consuming the (multiple) inputs starting from the top of rectangle 501a and ending at the bottom of the rectangle, producing output 504a. On the other hand, rectangle 502a represents the memory location changing over time. Thus, at the entry of the function, the value of the memory location is shown as "A", which is used as the input to the function. Arrows 505a and 505b represent some possible loads and stores to the memory location during the execution of the function (although there may be other loads and stores). As shown, at the store at arrow 505a, the value "B" is written to the memory location by something other than the function (e.g., another thread of the same process, a kernel thread, DMA, etc.). This value is later read by the function at arrow 505b. As is to be understood, the store at arrow 505a may be intentional (in which case the value "B" would be expected by the function at the load at arrow 505b), but it may also be due to a data race (in which case the value "A" would be expected by the function at the load at arrow 505b). Given this situation, the load value identification component 115d may choose to replace the value "A" into the load at arrow 505b. In this case, this may be an attempt to simulate the correction of a data race. It is noted that if there is no store at arrow 505a in the trace, the load value identification component 115d may choose to replace some other value into the load at arrow 505b to attempt to simulate the introduction of a data race.

[0064] Example 500b illustrates a load from a memory location that was previously an input. Specifically, rectangle 501b represents the execution of a function over time—consuming input(s) 503b starting at the top of the rectangle and ending at the bottom of the rectangle, producing output 504b. Rectangle 502b represents the memory location changing over time. Thus, at the entry of the function, the value of the memory location is shown as “A”, which is used as an input to the function. Arrows 505c through 505d represent some possible loads and stores to the memory location during the execution of the function (although there may be other loads and stores). As shown, at the store at arrow 505c, the function writes the value “B” to the memory location. Later, at arrow 505d, the value “C” is written to the memory location by something other than the function (e.g., another thread of the same process, a kernel thread, DMA, etc.). This value is later read by the function at arrow 505e. As is to be understood, the store at arrow 505d may be intentional (in which case the value “C” would be expected by the function at the load at arrow 505e), but it could also be due to a data race (in which case the value “B” would be expected by the function at the load at arrow 505e). Given this situation, the load value identification component 115d may choose to substitute the value “B” into the load at arrow 505e. In this case, this may be an attempt to simulate a correction of the data race. It is noted that if there is no store at arrow 505d in the trace, the load value identification component 115d may choose to substitute some other value into the load at arrow 505e to attempt to simulate the introduction of a data race.

[0065] Example 500c illustrates a load from a memory location that does not correspond to an input. Specifically, rectangle 501c represents the execution of a function over time—consuming the input(s) 503c (if any) starting from the top of the rectangle and ending at the bottom of the rectangle, producing output 504c. Rectangle 502c represents the memory location changing over time. Arrows 505f through 505h represent some possible loads and stores to the memory location during the execution of the function (although there may be other loads and stores). As shown, at the store at arrow 505f, the function writes the value “A” to the memory location. Later, at arrow 505g, the value “B” is written to the memory location by something other than the function (e.g., another thread of the same process, a kernel thread, DMA, etc.). This value is later read by the function at arrow 505h. As will be appreciated, the store at arrow 505g may be intentional (in which case the value “B” would be expected by the function at the load at arrow 505h), but it may also be due to a data race (in which case the value “A” would be expected by the function at the load at arrow 505h). Given this situation, the load value identification component 115d may choose to substitute the value “A” into the load at arrow 505h. In this case, this may be an attempt to simulate the correction of a data race. It is noted that if the store at arrow 505g does not exist in the trace, the load value identification component 115d may choose to substitute some other value into the load at arrow 505h to attempt to simulate the introduction of a data race.

[0066] Turning to the emulation component 116, the emulation component 116 emulates the code accessed by the code access component 114b based on one or more of the recorded executions 113 accessed by the trace access component 114a. For example, the emulation component 116 may include or utilize Figure 1A the emulation component 111 to emulate the accessed code. Using the emulation component 116, the debug component 109 may replay one or more functions identified by the function identification component 115a based on the code for the executed function while guiding the execution of the code with the traced data values from the recorded execution 113 (including, for example, one or more traced inputs identified by the input / output identification component 115b). Thus, the emulation component 116 is shown as including an emulation guidance component 116a that may supply the traced data values to the code of any emulated function as needed to guide the emulation of the function to reproduce the traced execution of the function.

[0067] The simulation component 116 is also shown as including a load replacement component 116b. In conjunction with the simulation of a given function, the load replacement component 116b can replace one or more of the values in a load identified by the load identification component 115c with values different from those that would be expected based on the operation of the simulation guidance component 116a. Thus, the load replacement component 116b can be considered to produce an exception to the guidance of the simulation guidance component 116a. The load replacement component 116b can produce replacement load values for any of the categories of loads discussed above in connection with the load identification component 115c. For example, if the load is from a location corresponding to an input, the load replacement component 116b can replace the input value obtained from the recorded execution 113 or obtained by replaying another function with some other value. If the load is from a location that previously corresponded to an input but was subsequently written to by a function, the load replacement component 116b can replace the written value with some other value. If the load is from a location that does not correspond to any input but after the function has performed one or more stores to that location, the load replacement component 116b can replace the written value with some other value.

[0068] The simulation component 116 is also shown as including an output generation component 116c. The output generation component 116c indicates that the code simulation will generate one or more output values during the code simulation. For example, the simulation of function 402-1 will produce output 404-1, the simulation of function 402-2 will produce output 404-2, etc. If the simulation component 116 utilizes the simulation guidance component 116a - rather than the load replacement component 116b - the simulation of a given function will produce the same outputs and output values as during tracing. Thus, the simulation guidance component 116a can be used to obtain any outputs that may not have been recorded during tracing. However, if the simulation component 116 also utilizes the load replacement component 116b, the simulation of the function may deviate from how the function executed during tracing, and it may produce different output values. Thus, for a given function, the output generation component 116c may be able to generate a set of one or more outputs and a set of one or more additional outputs, the (multiple) values of the one or more outputs being consistent with those output values produced during tracing, depending on which loads within the function the load replacement component 116b performs replacements on and / or which values the load replacement component 116b replaces, the set of one or more additional outputs may have different values.

[0069] Notably, when performing replay (including replay with memory load replacement), the simulation component 116 can utilize the tracker component 110 to record the simulation. Thus, the simulation component 116 can contribute additional traces to the recorded execution 113.

[0070] As mentioned, the function may include gaps, such as gaps caused by calls to kernel calls that are not traced. In some embodiments, the emulation component 116 may use one or more techniques to gracefully handle any gaps. As a first example, the emulation component 116 may determine from the accessed recorded execution 113 which inputs are supplied to the kernel call, and then the emulation component 116 may emulate the kernel call based on these inputs. As a second example, the emulation component 116 may treat the kernel call as an event that can be ordered among other events in the accessed recorded execution 113, and the emulation component 116 may ensure that any visible changes made by the kernel call (e.g., changed memory values, changed register values, etc.) are exposed as inputs to the code that executes after the kernel call, rather than emulating the kernel call. As a third example, the emulation component 116 may set the appropriate environment context, and then make an actual call to the running kernel using these inputs. As a fourth example, the emulation component may simply prompt the user for the result of the kernel call.

[0071] The emulation analysis component 117 may perform various types of analysis on the emulated execution of the accessed application(s) 112, including analyzing the impact of any load replacements performed by the load replacement component 116b. For example, this analysis may be used as part of a data race analysis, as will be discussed further. As shown, the emulation analysis component 117 may include, for example, an output comparison component 117a, a data race analysis component 117b, a classification component 117c, and / or an inspector component 117d.

[0072] The output comparison component 117a may compare the set of outputs generated by a given function during tracing with the set of outputs generated during replay of the function by the emulation component 116 when replacing the value(s) read by one or more function loads using the load replacement component 116b. If there are any differences within these sets of outputs, the output comparison component 117a may determine that the load replacement has an impact on the output of the function. As mentioned, the set of outputs generated during replay may be obtained from the recorded execution 113, and / or based on the emulation of the function by the emulation component 116 when the emulation bootstrapping component 116a uses the trace data to guide the emulation.

[0073] Based on the comparison by the output comparison component 117a, the data race analysis component 117b can determine whether the load replacement component 116b, which has replaced a value for a given load, has possibly corrected a data race or has possibly simulated a data race. If the output comparison component 117a finds no difference between the set of traced outputs and the set of outputs generated during the load replacement, the replacement may not have corrected or simulated the data race (at least not in a detectable way). However, if there is a difference in the output sets, the replacement may have corrected or simulated the data race. In some embodiments, the data race analysis component 117b can determine whether the simulated output set has normal values or outliers. If the correction of the race is simulated and if the resulting output is normal, there may indeed be a race. Conversely, if the correction of the race is simulated and if the resulting output is abnormal, there may be no race, or the values used to simulate the correction are a poor choice. Similarly, if the introduction of the race is simulated, observing whether the output is normal or abnormal can indicate whether the simulated race matches what might occur in production. For example, if the output remains normal, the simulated race may not actually occur in production. Conversely, if the output is abnormal and if they are similar to the abnormal results observed in production, the simulated race may be occurring in production.

[0074] To determine whether the simulated output is normal or abnormal, the simulation analysis component 117 can include a classification component 117c. The classification component 117c can take as input the outputs of multiple traced instances of a given function. The classification component 117c can then look for patterns in the outputs (e.g., using machine learning techniques) to classify the various outputs as normal or abnormal for that function. For example, the classification component 117c can determine which output values typically fall within a normal distribution with respect to each other (and are therefore normal), and / or which output values fall outside of that distribution (and are therefore abnormal). Regardless of the particular classification technique used, the classification component 117c can develop a model for a given function that can be used to determine which output values for that function are normal or abnormal. In some embodiments, the classification component 117c can also analyze the inputs to the function and model the outputs typical for a given input.

[0075] The simulation analysis component 117 may also include an inspector component 117d that can perform one or more queries on the recorded executions 113—whether the recorded executions 113 were generated based on “live” code execution or whether they were generated in conjunction with the simulation. These queries can examine various types of code behavior, such as memory leaks (e.g., by querying for any memory allocations that do not have a corresponding deallocation). In some embodiments, the simulation analysis component 117 may run one or more inspectors against a trace recorded based on the execution of the memory load replacement by the simulation component 116. Thus, the simulation analysis component 117 can determine whether the memory load replacement causes (or stops) behavior detectable by the inspector component 117d. The data race analysis component 117b can use the inspector component 117d to determine whether the memory load replacement may have corrected a data race or may have mimicked a part of a data race. For example, if an undesired behavior detectable by the inspector component 117d is fixed after the memory load replacement, this may be an indication that the memory load replacement fixed a memory race. Conversely, if an undesired behavior detectable by the inspector component 117d is introduced after the memory load replacement, this may be an indication that the memory load replacement caused a memory race.

[0076] The output component 118 may output the results of any code simulation by the simulation component 116 and / or the results of any analysis by the simulation analysis component 117. For example, the output component 118 may visualize the code simulation (including load value replacement), may generate a notification (e.g., to another application or at a user interface) in the case where the output comparison component 117a determines that the simulated output differs from the traced output, may present any differences between the traced output and the simulated output, etc.

[0077] In view of the foregoing, Figure 6 A flowchart of an example method 600 for using recorded executions to determine whether a load modification inside a function affects one or more outputs of the traced function is illustrated. Method 600 will now be described in the context of FIGS. 1 through 5C. Although, for purposes of illustration, the acts of method 600 are shown in a particular sequential order, it is to be appreciated that some of these acts may be implemented in a different order and / or in parallel.

[0078] As Figure 6As shown, method 600 includes an action 601 of accessing (multiple) previously executed (multiple) replayable traces of an executable entity. In some embodiments, action 601 includes accessing a recorded execution that includes trace data of a previous execution of at least a portion of the executable code of the executable entity, the trace data enabling the replay of the previous execution of that portion of the executable entity. For example, data access component 114 may access one or more recorded executions 113 of application 112 (e.g., using trace access component 114a). As Figure 3 shown, each of these recorded executions 113 may include at least one data stream 301a, and data stream 301a includes multiple data packets 302; each data packet 302 may include a data input portion 304 that records the input to an executable instruction that was executed as part of a previous execution of the application. The (multiple) recorded executions 113 may include a previous "live" execution of application 112 at (multiple) processors 102 directly or through a managed runtime or a previously simulated execution of application 112 using simulation component 116. Thus, in action 601, one or more recorded executions 113 may include at least one of a "live" execution of the executable entity or a simulated execution of the executable entity.

[0079] Method 600 further includes an action 602 of identifying functions within the executable entity. In some embodiments, action 602 includes identifying functions within a traced portion of the executable code of the executable entity, the functions including sequences of executable instructions that consume zero or more inputs and produce one or more outputs. For example, function identification component 115a may analyze the code of application 112 (e.g., included in the application binary and / or the trace data stream) to identify one or more functions in the code. As discussed, using the definitions of inputs and outputs herein, a function is a sequence of executable instructions having zero or more inputs and one or more outputs.

[0080] Method 600 also includes an action 603 of determining the traced output of a function. In some embodiments, action 603 includes determining, based on the trace data, one or more traced output data values generated by the traced instance of the function during a previous execution. For example, the input / output identification component 115b may identify one or more outputs of the function identified in action 602, including the (multiple) output values of the subject instance of the function. The output values may be obtained directly from the recorded execution 113, or obtained based on replaying the function using the simulation component 116 (in conjunction with the simulation guidance component 116a). Although not shown, method 600 may also use the input / output identification component 115b to determine one or more traced input data values (if any) provided to the traced instance of the function during a previous execution.

[0081] Method 600 also includes an action 604 of identifying internal loads within the function. In some embodiments, action 604 includes identifying at least one executable instruction that performs a load from memory within the sequence of executable instructions of the function. For example, the load identification component 115c may analyze the executable instructions of the function to identify one or more loads within the function. For example, these loads may target system memory, registers, etc. In some embodiments, the load identification component 115c specifically identifies the (multiple) executable instructions that perform loads from memory that are likely to experience memory contention. For example, these loads may target heap memory, global variables, or other memory locations that may be accessed by multiple threads. In some embodiments, the load identification component 115c only identifies those loads that may affect the (multiple) output values identified in action 603. Thus, the load identification component 115c may perform a "backward contamination" analysis of the sequence of executable instructions to identify loads that have data continuity with the (multiple) outputs.

[0082] Method 600 also includes an action 605 of simulating the execution of the function according to the trace. For example, the simulation component 116 may simulate the sequence of executable instructions of the function while using the simulation guidance component 116a together with the trace data included in the accessed recorded execution 113 to guide their execution. Thus, the simulation component 116 may reproduce the previous execution of the function as recorded in the recorded execution 113. Although not shown, action 605 may include providing one or more traced input data values for the function for one or more inputs (if any) when simulating the execution of the function according to the trace data.

[0083] Although the use of the emulation guidance component 116a can replay a previous execution, the emulation component 116 can also modify the replay by replacing the loaded traced values with different values (e.g., the values identified by the load value identifying component 115d). Thus, as shown, action 605 can include an action 605a of replacing the loaded memory values. In some embodiments, action 605a includes replacing the traced memory values loaded by at least one executable instruction during a traced instance of a function with different memory values. For example, when emulating the load identified in action 604, the load replacement component 116b can replace the value that would be used based on the recorded execution 113 with a different value identified by the load value identifying component 115d. For example, this can include replacing the traced memory value with a memory value from one or more traced input data values (e.g., to simulate a corrected data race), or replacing the traced memory value with some other memory value. As shown, action 605 can include an action 605b of generating an emulated output from the function. Thus, as represented by the output generation component 116c, the emulation component 116 can generate one or more emulated output data values for one or more outputs.

[0084] As will be appreciated, due to the use of the load replacement component 116b, these outputs may be different from the outputs that would be generated by a standard replay of the function. Thus, method 600 includes an action 606 of comparing the traced output and the emulated output. In some embodiments, action 606 includes determining whether there is a difference between one or more traced output data values and one or more emulated output data values. For example, the output comparison component 117a can compare the output values obtained in action 603 with the output values generated in action 605b. Then, method 600 further includes an action 607 of generating a result from the comparison. The result may be that the outputs are the same, or the outputs are different. If the outputs are different, action 607 can include: based on the existence of a difference between one or more traced output data values and one or more emulated output data values, generating a notification at the user interface or to a software component. If a notification is generated to a software component (which can include, for example, calling a function), then action 607 can include determining whether one or more emulated output data values are normal or abnormal for the function, determining whether one or more checkers are affected by replacing the traced memory values with different memory values, etc. For example, the output component 118c can generate an alert regarding the difference and / or the result of any analysis performed by the data race analysis component 117b. The data race analysis component 117b can determine whether the emulated output is normal or abnormal (e.g., using the classification component 117c), and the data race analysis component 117b can determine whether the load replacement causes a change in the checker (e.g., using the checker component 117d), etc.

[0085] If, in operation 607, the data race analysis component 117b determines whether one or more of the output data values of the simulation are normal or abnormal for the function, the determination can compare one or more of the output data values of the simulation with the traced output values from one or more other traced instances of the function, such as by using the classification component 117c. If, in operation 607, the data race analysis component 117b determines that one or more of the output data values of the simulation are normal for the function, method 600 can include the data race analysis component 117b determining to replace the traced memory values with different memory values that correct potential race conditions. Alternatively, if, in operation 607, the data race analysis component 117b determines that one or more of the output data values of the simulation are abnormal for the function, method 600 can include the data race analysis component 117b determining to replace the traced memory values with different memory values that introduce potential race conditions. If, in operation 607, the output component 118 generates a notification, the notification can cause the user interface to present one or more of the traced output data values and one or more of the output data values of the simulation.

[0086] Accordingly, embodiments herein utilize historical debugging techniques to replace loaded values within a function during function replay to observe the effect of the replacement on the function output(s). Modifying the value read by a given load can be manipulated to simulate the occurrence of a data race. Thus, the debugger can observe the effect that a lock-simulated data race might have on the function behavior, even when no data race actually occurs during tracing. Accordingly, when a race can be suspected but not captured during tracing, the debugger can implement data race simulation within the function. Alternatively, modifying the value read by a given load can be manipulated to simulate the correction of a data race. Thus, when a data race is captured during tracing, the debugger can observe the effect that a simulated correction of the data race might have on the behavior of the function. This can reduce the time spent on development, troubleshooting, and debugging, resulting in the production of less error-prone code.

[0087] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts or the order of the acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0088] The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. Thus, the scope of the present invention is indicated by the appended claims rather than by the foregoing description. All changes that fall within the equivalent meaning and scope of the claims are to be embraced within their scope. When an element is introduced in the appended claims, the articles "a", "an", "the" and "said" are intended to mean that there is one or more of the element. The terms "comprising", "including" and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.

Claims

1. A method implemented at a computer system including one or more processors and a memory for determining whether a load modification within a function affects one or more outputs of a traced function using a recorded execution, the method comprising: Accessing the recorded execution, the recorded execution including trace data that traces a previous execution of at least a portion of the executable code of an executable entity, the trace data enabling the replay of the previous execution of the portion of the executable entity; Identifying a function within the traced portion of the executable code of the executable entity, the function including a sequence of executable instructions, the function consuming zero or more inputs and producing one or more outputs; Based on the trace data, determining one or more traced output data values produced by a traced instance of the function during the previous execution; Identifying at least one executable instruction that performs a load from memory within the sequence of executable instructions of the function; Simulating the execution of the function according to the trace data, including: Replacing a traced memory value loaded by the at least one executable instruction during the traced instance of the function with a different memory value; and Producing one or more simulated output data values for the one or more outputs; Comparing the one or more traced output data values and the one or more simulated output data values, including determining whether there is a difference between the one or more traced output data values and the one or more simulated output data values; As part of a data race analysis, analyzing the impact of replacing the traced memory value with a different memory value, including determining based on the comparison whether replacing the traced memory value with the different memory value corresponds to a correction of a data race or a simulation of a data race; and Outputting the result of the simulation of the function and the result of the analysis, including generating a notification at a user interface or to a software component based on a difference between the one or more traced output data values and the one or more simulated output data values.

2. The method according to claim 1, wherein identifying the at least one executable instruction that performs the load from the memory comprises: Identifying executable instructions that perform a load from memory that can be subject to a memory race.

3. The method according to claim 1, wherein identifying the at least one executable instruction that performs the load from the memory comprises: Based on an analysis of the sequence of executable instructions, identifying values loaded that can be traced to values of the one or more outputs.

4. The method according to claim 1, wherein the method generates the notification to the software component, and wherein analyzing the impact of replacing the traced memory value with a different memory value includes determining whether the one or more simulated output data values are normal or abnormal for the function, wherein the determining includes classifying the one or more simulated output data values as normal or abnormal for the function based on a pattern in traced output values from one or more other traced instances of the function.

5. The method according to claim 4, wherein analyzing the effect of replacing the traced memory value with different memory values includes: Based on determining that the output data values of the one or more simulations are normal for the function and determining that replacing the traced memory value with the different memory values simulates a potential race condition correction, it is determined that a data race may occur.

6. The method according to claim 1, wherein said analyzing the effect of replacing the traced memory value with different memory values comprises: Based on determining that the output data values of the one or more simulations are abnormal for the function and determining that replacing the traced memory value with the different memory values simulates the introduction of a potential race condition, it is determined that a data race may occur.

7. The method according to claim 1, wherein the method generates the notification at the user interface, and the notification causes the user interface to present the one or more traced output data values and the output data values of the one or more simulations.

8. The method according to claim 1, wherein the method generates the notification to the software component, and wherein the method further includes performing one or more queries on the recorded execution, the queries being configured to examine various types of code behavior and determining whether one or more checkers are affected by replacing the traced memory value with the different memory values.

9. The method according to claim 1, wherein replacing the traced memory value with the different memory values comprises: Replace the traced memory value with a memory value from one or more traced input data values.

10. The method according to claim 1, further comprising: Based on the trace data, determining one or more traced input data values for one or more inputs provided to the traced instance of the function during the previous execution; and When simulating the execution of the function according to the trace data, providing the one or more traced input data values for the one or more inputs to the function.

11. A computer system, comprising: One or more processors; and One or more computer-readable media having computer-executable instructions stored thereon, the computer-executable instructions being executable by the one or more processors to cause the computer system to use the recorded execution to determine whether a load modification within a function affects one or more outputs of the traced function, the computer-executable instructions including instructions executable by the one or more processors to at least perform the following: Access the recorded execution, the recorded execution including trace data that traces a previous execution of at least a portion of the executable code of an executable entity, the trace data enabling the replay of the previous execution of the portion of the executable entity; Identify a function within the traced portion of the executable code of the executable entity, the function including a sequence of executable instructions, the function consuming zero or more inputs and producing one or more outputs; Based on the trace data, determine one or more traced output data values produced by the traced instance of the function during the previous execution; Identify at least one executable instruction that performs a load from memory within the sequence of executable instructions of the function; Simulate the execution of the function according to the trace data, including: Replace the traced memory values loaded by the at least one executable instruction during the traced instance of the function with different memory values; and Generate one or more simulated output data values for the one or more outputs; Compare the one or more traced output data values and the one or more simulated output data values, including determining whether there are differences between the one or more traced output data values and the one or more simulated output data values; As part of the data race analysis, analyze the impact of replacing the traced memory values with different memory values, including determining whether replacing the traced memory values with the different memory values corresponds to a correction of a data race or a simulation of a data race based on the comparison; and Output the results of the simulation of the function and the results of the analysis, including generating a notification at a user interface or to a software component based on there being differences between the one or more traced output data values and the one or more simulated output data values.

12. The computer system according to claim 11, wherein identifying the at least one executable instruction that performs the load from the memory includes: Identify executable instructions that perform loads from memory that can be subject to memory contention.

13. The computer system according to claim 11, wherein identifying the at least one executable instruction that performs the load from the memory includes: Based on an analysis of the sequence of executable instructions, identify values loaded that can be traced to values of the one or more outputs.

14. The computer system of claim 11, wherein the computer system generates the notification to the software component, and wherein analyzing the impact of replacing the traced memory values with different memory values includes determining whether the one or more simulated output data values are normal or abnormal for the function, wherein the determining includes classifying the one or more simulated output data values as normal or abnormal for the function based on patterns in traced output values from one or more other traced instances of the function.

15. The computer system according to claim 11, wherein the analysis of the effect of replacing the traced memory value with a different memory value includes: Based on determining that the one or more simulated output data values are normal for the function and determining at least one of the following, determine that a data race may occur: Replacing the traced memory values with the different memory values simulates a correction of a potential contention condition; or Replacing the traced memory values with the different memory values simulates the introduction of a potential contention condition.

16. The computer system of claim 11, wherein the computer system generates the notification at the user interface, and the notification causes the user interface to present the one or more traced output data values and the one or more simulated output data values.

17. The computer system of claim 11, wherein the computer system generates the notification to the software component, and the computer system also performs one or more queries on the recorded execution, the queries being configured to examine various types of code behavior and determine whether one or more checkers are affected by replacing the traced memory values with the different memory values.

18. The computer system according to claim 11, wherein replacing the traced memory value with the different memory values comprises: Replace the traced memory values with memory values from one or more traced input data values.

19. The computer system according to claim 11, wherein the computer-executable instructions further comprise instructions executable by the one or more processors to perform the following: Based on the trace data, determine one or more traced input data values for one or more inputs provided to the traced instances of the function during the previous execution; and When simulating the execution of the function according to the trace data, provide the one or more traced input data values for the one or more inputs to the function.

20. A computer program product comprising one or more hardware storage devices having stored thereon computer-executable instructions executable by one or more processors to cause a computer system to use a recorded execution to determine whether a load modification within a function affects one or more outputs of a traced function, the computer-executable instructions comprising instructions executable by the one or more processors to at least perform the following: Access the recorded execution, the recorded execution including trace data that traces a previous execution of at least a portion of the executable code of an executable entity, the trace data enabling the replay of the previous execution of the portion of the executable entity; Identify a function within the traced portion of the executable code of the executable entity, the function including a sequence of executable instructions, the function consuming zero or more inputs and producing one or more outputs; Based on the trace data, determine one or more traced output data values produced by the traced instances of the function during the previous execution; Identify at least one executable instruction that performs a load from memory within the sequence of executable instructions of the function; Simulate the execution of the function according to the trace data, including: Replacing the traced memory value loaded by the at least one executable instruction during the traced instance of the function with a different memory value; and Producing one or more simulated output data values for the one or more outputs; Compare the one or more traced output data values and the one or more simulated output data values, including determining whether there is a difference between the one or more traced output data values and the one or more simulated output data values; As part of a data race analysis, analyze the effect of replacing the traced memory value with a different memory value, including determining based on the comparison whether replacing the traced memory value with the different memory value corresponds to a correction of a data race or a simulation of a data race; and Output the result of the simulation of the function and the result of the analysis, including generating a notification at a user interface or to a software component based on a difference between the one or more traced output data values and the one or more simulated output data values.

Citation Information

Patent Citations

  • Tentative execution of code in a debugger

    US20190042396A1