Using entropy to prevent payload data from being included in code execution log data

JP2024521135A5Active Publication Date: 2025-05-13MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023572175
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-21
Filing Date
2022-05-02
Publication Date
2025-05-13
Estimated Expiration
2042-05-02

AI Technical Summary

Technical Problem

Existing diagnostic tools and techniques for software debugging often inadvertently include sensitive data such as personally identifiable information (PII), cryptographic keys, and passwords in replayable execution traces and execution logs, posing a security risk and consuming excessive computing resources.

Method used

Employ entropy analysis to identify high-entropy data items, such as PII and cryptographic keys, and exclude them from code execution logs by replacing them with constraints or alternative data that preserve code flow, thereby preventing their inclusion and reducing log data size.

Benefits of technology

Enhances data security by preventing sensitive data exposure and reducing the size of execution logs, conserving computing resources by minimizing the storage and processing required for log analysis and transfer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Using entropy to prevent inclusion of payload data in code execution log data. Some embodiments determine that a payload data item associated with code execution log data has entropy above a predetermined entropy threshold and identify the particular executable code that interacted with the payload data item. Some embodiments then take preventive action to exclude the payload data item from being included in execution records of the particular executable code. Examples of preventive action include preventing the payload data item from being exported from a computer system, preventing the payload data item from being included in code execution log data, and adding the payload data item to a block list with reference to the particular executable code.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to systems, methods, and devices that prevent sensitive payload data, such as personally identifiable information (PII), encryption keys, and passwords, from being included in replayable execution traces and other execution logs. [Background technology]

[0002]

[0002] Tracking and correcting undesirable software behaviors / faults is a core activity of software development. Undesirable software behaviors can include many things, such as execution crashes, runtime exceptions, slow execution performance, incorrect data results, data corruption, etc. Undesirable software behaviors are triggered by various factors, such as data input, user input, race conditions (e.g., when accessing a shared resource), etc. Given the various triggers, undesirable software behaviors are often rare, seemingly random, and very difficult to reproduce. Therefore, it is often very time-consuming and difficult for developers to identify a given undesirable software behavior. Even if an undesirable software behavior is identified, it is often again time-consuming and difficult to determine its root cause.

[0003]

[0003] Developers use various techniques to identify undesirable software behavior and then identify one or more locations in the application's code that cause the undesirable software behavior. For example, developers often test various portions of the application's code against various inputs (e.g., unit tests). As another example, developers often reason about the execution of the application's code in a debugger (e.g., by setting breakpoints / watchpoints, stepping through code lines as the code is executed, etc.). As another example, developers often observe code execution behavior (e.g., timing, coverage, etc.) with a profiler. As another example, developers often insert diagnostic code (e.g., trace statements) into the application's code. Each of these activities is aided by code execution logs, such as event logs generated by diagnostic code included in the target application.

[0004]

[0004] While traditional diagnostic tools (e.g., debuggers, profilers, etc.) operate on "live" forward-executing code, new forms of diagnostic tools enable "historical" debugging (also called "time travel" or "reverse" debugging), where at least a portion of the execution context (e.g., processes, threads, etc.) of an executable computer program is recorded in a code execution log that includes one or more trace files (i.e., execution traces). Using some tracing techniques, the execution traces include "bit-accurate" historical execution trace data. This allows any recorded portion of the traced execution context to be "replayed" virtually (e.g., via emulation) down to the granularity of individual instructions (e.g., machine code instructions, intermediate language code instructions, etc.). Thus, using the bit-accurate trace data, diagnostic tools allow developers to reason about the recorded previous executions of the executable program, in contrast to traditional debugging, which is limited to "live" forward execution. For example, using replayable execution traces, some historical debuggers provide a user experience that allows, for example, both forward and reverse breakpoints / watchpoints, or allows stepping through code both forward and reverse. On the other hand, some historical profilers can derive code execution behavior (e.g. timing and coverage) from previously executed code.

[0005]

[0005] A replayable execution trace explicitly or implicitly includes every input to and every output from each recorded instruction. Thus, a replayable execution trace includes all data consumed or generated by the traced code. When the traced code consumes or generates sensitive data items, such as PII, cryptographic keys, passwords, etc., the execution trace of the traced code also includes these sensitive data items. Less strict forms of code execution logs, such as event logs generated by diagnostic code in an application, may also include sensitive data items. Summary of the Invention

[0006] At least some embodiments described herein improve data security by using entropy analysis to prevent inclusion of payload data in code execution log data. These embodiments address the technical challenge of being able to efficiently and reliably determine that a particular data item should be considered a sensitive data item (e.g., PII, cryptographic keys, passwords, etc.). These embodiments are based on the inventor's observation that sensitive data items tend to have relatively high entropy when compared to less sensitive data items (e.g., environmental variables, mathematical constants, etc.). In some embodiments, the entropy of a data item is considered intrinsically (i.e., by looking at the ratio of the number of bits of entropy to the total length of the data item) and / or contextually (i.e., by determining the uniqueness of the value of a given data item in one code execution log dataset compared to the values ​​of the data item in other code execution log datasets). Recognizing that sensitive data items have relatively high entropy, these embodiments identify high entropy data items and exclude them from inclusion in code execution log data such as replayable execution traces and event logs.

[0007]

[0007] In some embodiments, excluding high entropy data items from the code execution log data has the technical effect of promoting data security by preventing inadvertent exposure of sensitive data items via the code execution log data. Furthermore, excluding high entropy data items from the code execution log data reduces the size of the code execution log data by removing portions of the data, including removing high entropy data that is often not easily compressible, which has the additional technical effect of conserving computing resources. For example, reducing the size of the code execution log data saves processing resources when analyzing that log data and saves storage and network resources required to store and transfer the code execution log data.

[0008] In accordance with the above, some embodiments are directed to using entropy to prevent inclusion of payload data in code execution log data. These embodiments determine that a payload data item associated with the code execution log data has entropy that exceeds a predetermined entropy threshold. Based on a determination that the payload data item has entropy that exceeds the predetermined entropy threshold, these embodiments identify a particular executable code that interacted with the payload data item. These embodiments then take preventative action to exclude the payload data item from being included in the execution record of the particular executable code.

[0009] At least some additional or alternative embodiments described herein improve data security by removing up to all payload data from the execution trace without compromising the ability to perform some forms of execution trace analysis that do not rely on analysis of the payload data itself (e.g., code and / or data flow analysis, memory locality analysis, memory access pattern analysis, cache usage analysis, race condition analysis, buffer overflow analysis, traditional debugging, etc.). For example, some embodiments process the execution trace to identify payload data items. For each identified payload data item, some embodiments determine constraints that execution of code that interacted with the data item imposed on the data item, and replace the value of the data item in the execution trace with information that maintains those constraints. In some embodiments, this information includes the executable code itself, memory addresses of the data items, and / or data structured to preserve the code flow (e.g., replacement values ​​for the data items, specifications of valid values ​​for the data items, instructions of the code path to follow, etc.).

[0010]

[0010] In some embodiments, excluding payload data from execution traces has the technical effect of promoting data security by preventing inadvertent exposure of sensitive data items via the execution traces. Moreover, in some embodiments, excluding payload data from execution traces can significantly increase the compression ratio of those execution traces. In any event, excluding payload data from execution traces can significantly reduce the size of the execution traces, which has the additional technical effect of saving computing resources. For example, reducing the size of the execution traces saves processing resources when analyzing those execution traces and saves storage and network resources required when storing and transferring the execution traces.

[0011] In accordance with the above, in some embodiments, methods, systems, and computer products are directed to removing payload data from an execution trace. These embodiments identify a payload data item in the execution trace and identify a particular executable code that interacted with the payload data item. Based on the payload data item and the particular executable code, these embodiments determine one or more constraints that execution of the particular executable code imposes on the payload data item, and then replace the value of the payload data item in the execution trace with information that maintains the one or more constraints. In various embodiments, the one or more constraints include one or more of bytes of the particular executable code, memory addresses corresponding to the payload data item, or data structured to preserve code flow, and the information that maintains the one or more constraints includes one or more of bytes of the particular executable code, memory addresses corresponding to the payload data item, or data structured to preserve code flow.

[0012]

[0012] This Summary is provided to introduce some concepts in a simplified form that are further described below in the Detailed Description section. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0013]

[0013] In order to set forth how the above and other advantages and features of the invention can be obtained, a more particular description of the invention briefly described above will be rendered by reference to specific embodiments thereof which are illustrated in the accompanying drawings, in which the invention will be described and explained with further specificity and detail, with the understanding that these drawings depict only typical embodiments of the invention and therefore should not be considered as limiting the scope of the invention. [Brief description of the drawings]

[0014] [Figure 1]

[0014] FIG. 1 illustrates an exemplary computing environment that facilitates one or more of using entropy to prevent payload data from being included in code execution log data or removing payload data from an execution trace. [Diagram 2]

[0015] FIG. 13 illustrates additional details of a debug component configured to use entropy to prevent payload data from being included in code execution log data. [Diagram 3]

[0016] FIG. 2 illustrates details of a debug component configured to remove payload data from an execution trace. [Figure 4]

[0017] FIG. 2 illustrates an exemplary computing environment in which the computer system of FIG. 1 is connected to one or more other computer systems via one or more networks. [Diagram 5]

[0018] FIG. 13 illustrates an example of an execution trace. [Figure 6]

[0019] 1 is a flow diagram of an example method for using entropy to prevent inclusion of payload data in code execution log data. [Figure 7]

[0020] FIG. 2 illustrates a flow diagram of an example method for removing payload data from an execution trace. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015]

[0021] As mentioned, at least some embodiments described herein are directed to one or more of: (i) embodiments for using entropy to prevent inclusion of payload data in code execution log data (such as execution traces or event logs), or (ii) embodiments for removing payload data from execution traces, either of which promote data security and reduce the size of code execution log data / execution traces while addressing the technical challenge of efficiently and reliably determining that certain data items should be considered sensitive data items.

[0016]

[0022] In some embodiments, the execution traces used by the embodiments herein are generated by one or more of a variety of historical debugging techniques. In general, historical debugging techniques record or reconstruct the state of an execution context (e.g., a process, thread, etc.) at various points in time as the code of a corresponding executable computer program is executed, allowing that execution context to be at least partially replayed from that execution state. The fidelity of the virtual execution varies depending on the available traced execution states.

[0017]

[0023] In one example, some classes of historical debugging techniques, referred to herein as time-travel debugging, continuously record a bit-precise trace of an execution context. This bit-precise trace can later be used to replay previous executions of the execution context with fidelity down to the fidelity of individual code instructions. For example, the bit-precise trace records (e.g., by recording snapshots of processor registers) sufficient information to recreate an initial processor state for at least one point in time in a previous execution of the execution context, along with data values ​​(e.g., memory read values) read by the executable instructions when executed from that point onward. This bit-precise trace can then be used to replay the execution of those executable instructions (starting from the initial processor state) based on providing the recorded read values ​​to the instructions.

[0018]

[0024] Another class of historical debugging techniques, referred to herein as branch trace debugging, rely on reconstructing at least a portion of the state of an execution context based on reverse computation from a dump or snapshot (e.g., a crash dump) that includes a processor branch trace (i.e., includes a record of whether a branch was taken or not). These techniques start with values ​​(e.g., memory and registers) from this dump or snapshot, use the branch trace to at least partially determine the code execution flow, and repeatedly replay code instructions that were executed as part of the execution context forward and backward to reconstruct intermediate data values ​​(e.g., registers and memory) used by the code instructions until their values ​​reach a steady state. These techniques may be limited in how far back they can reconstruct data values ​​and how many data values ​​they can reconstruct. Nevertheless, the reconstructed historical execution data can be used for historical debugging.

[0019]

[0025] Yet another class of historical debugging techniques, referred to herein as replay and snapshot debugging, periodically records complete snapshots of an execution context's memory space and processor registers during execution. If the execution context relies on data from sources other than the execution context's own memory or from non-deterministic sources, in some embodiments these techniques also record such data along with the snapshots. These techniques then use the data in the snapshots to replay the execution of the executable program's code between snapshots.

[0020]

[0026] 1 illustrates an exemplary computing environment 100 that facilitates one or more of using entropy to prevent inclusion of payload data in code execution log data or removing payload data from execution traces. As shown, computing environment 100 includes a computer system 101 (e.g., a special-purpose or general-purpose computing device) that includes a processor 102 (or multiple processors). As shown, in addition to the processor 102, computer system 101 also includes system memory 103, persistent storage 104, and possibly a network device 105 (or multiple network devices), which are communicatively coupled to each other and to the processor 102 using at least one communication bus 106.

[0021]

[0027] Embodiments within the scope of the present invention may include physical media and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media may be any available media that can be accessed by a general-purpose or special-purpose computer system. Computer-readable media that store computer-executable instructions and / or data structures are computer storage media. Computer-readable media that carry computer-executable instructions and / or data structures are transmission media. Thus, by way of example, and not limitation, embodiments of the present invention may comprise at least two distinctly different kinds of computer-readable media: computer storage media and transmission media.

[0022]

[0028] A computer storage medium is a physical storage medium (e.g., system memory 103 and / or persistent storage 104) that stores computer-executable instructions and / or data structures. Physical storage media includes computer hardware such as RAM, ROM, EEPROM, solid-state drives ("SSD"), flash memory, phase-change memory ("PCM"), optical disk storage, magnetic disk storage or other magnetic storage devices, or any other hardware storage device that can be used to store program code in the form of computer-executable instructions or data structures that can be accessed and executed by a general-purpose or special-purpose computer system to implement the disclosed functions of the present invention.

[0023]

[0029] A transmission medium may include a network and / or data link, which may be used to carry program code in the form of computer-executable instructions or data structures and may be accessed by a general-purpose or special-purpose computer. A "network" is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided to a computer system over a network (e.g., network device 105) or another communications connection (either wired, wireless, or a combination of wired and wireless), the computer system may consider the connection a transmission medium. Combinations of the above should also be included within the scope of computer-readable media.

[0024]

[0030] Furthermore, upon reaching various computer system components, program code in the form of computer-executable instructions or data structures may be automatically transferred from transmission media to computer storage media (or vice versa). For example, computer-executable instructions or data structures received over a network or data link may be buffered in RAM in a network interface module (not shown) and then eventually transferred to the computer system's RAM (e.g., system memory 103) and / or to a low-volatility computer storage medium (e.g., persistent storage 104) in the computer system. It should thus be understood that computer storage media may be included in computer system components that also (or primarily) utilize transmission media.

[0025]

[0031] Computer-executable instructions include, for example, instructions and data that, when executed by one or more processors, cause a general-purpose computer system, special-purpose computer system, or special-purpose processing device to perform a certain function or group of functions. Computer-executable instructions may be, for example, machine code instructions (e.g., binaries), intermediate format instructions such as assembly language, or even source code.

[0026]

[0032] Those skilled in the art will appreciate that the present invention may be practiced in a network computing environment with many types of computer system configurations including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, cell phones, PDAs, tablets, pagers, routers, switches, and the like. The present invention may also be practiced in a distributed system environment where tasks are both performed by local and remote computer systems that are linked over a network (either by wired data links, wireless data links, or a combination of wired and wireless data links). Thus, in a distributed system environment, a computer system may include multiple constituent computer systems. In a distributed system environment, program modules may be located in both local and remote memory storage devices.

[0027]

[0033] Those skilled in the art will also appreciate that the present invention may be implemented in a cloud computing environment. A cloud computing environment may be distributed, but this is not required. If distributed, a cloud computing environment may have components that are distributed internationally within an organization and / or owned across multiple organizations. In this specification and in the claims that follow, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services). The definition of "cloud computing" is not limited to any of the many other benefits that can be derived from such a model when properly deployed.

[0028]

[0034] Cloud computing models can consist of various characteristics such as on-demand self-service, broad network access, resource pooling, rapid elasticity, scalable services, etc. Cloud computing models may be in the form of various service models such as, for example, Software as a Service ("SaaS"), Platform as a Service ("PaaS"), and Infrastructure as a Service ("IaaS"). Cloud computing models can also be deployed using various deployment models such as private cloud, community cloud, public cloud, hybrid cloud, etc.

[0029]

[0035] Some embodiments, such as a cloud computing environment, may comprise a system including one or more hosts, each capable of running one or more virtual machines. In operation, the virtual machines emulate a running computing system and support an operating system and possibly one or more other applications. In some embodiments, each host includes a hypervisor that emulates virtual resources for the virtual machines using physical resources abstracted from the perspective of the virtual machines. The hypervisor also provides appropriate isolation between the virtual machines. Thus, the hypervisor provides the illusion that the virtual machines are interfacing with physical resources, even though from the perspective of any given virtual machine, the virtual machine is only interfacing with the appearance of physical resources (e.g., virtual resources). Examples of physical resources include processing power, memory, disk space, network bandwidth, media drives, etc.

[0030]

[0036] As shown in FIG. 1, in some embodiments, the processor 102 includes, for example, a processing unit 107 (or multiple processing units) and a memory cache 108 (or multiple memory caches). Each processing unit 107 (e.g., a processor core) loads and executes machine code instructions into at least one of multiple execution units 107b. During execution of these machine code instructions, the instructions may use registers 107a as temporary storage locations and may read and write to various locations in the system memory 103 via the memory cache 108. Each processing unit 107 executes machine code instructions that are defined by a processor instruction set architecture (ISA). The particular ISA of each processor 102 may vary based on the processor manufacturer and processor model. Common ISAs include the IA-64 and IA-32 architectures from INTEL, INC., the AMD64 architecture from ADVANCED MICRO DEVICES, INC., and various Advanced RISC Machine ("ARM") architectures from ARM HOLDINGS, PLC, although many other ISAs exist and may be used by the present invention. As is commonly understood, a machine code instruction is the smallest externally visible (ie, outside the processor) unit of code that can be executed by a processor.

[0031]

[0037] Registers 107a are hardware storage locations defined based on the ISA of processor 102. Registers 107a are read from and / or written to machine code instructions or processing units 107 as those instructions are executed in execution units 107b. Registers 107a are generally used to store values ​​fetched from memory cache 108 for use as inputs for executing machine code instructions, to store results of executing machine code instructions, to store program instruction counts, to support thread stack maintenance, etc. In some embodiments, registers 107a include "flags" that are used to signal any state changes caused by executing machine code instructions (e.g., to indicate whether an arithmetic operation was a carry, a zero result, etc.). In some embodiments, registers 107a include one or more control registers (e.g., used to control different aspects of processor operation) and / or other processor model specific registers (MSRs).

[0032]

[0038] The memory cache 108 temporarily caches blocks of the system memory 103 during execution of machine code instructions by the processing unit 107. In some implementations, the memory cache 108 includes one or more "code" portions that cache portions of the system memory 103 that store application code, and one or more "data" portions that cache portions of the system memory 103 that store application runtime data. If the processing unit 107 requests data (e.g., code or application runtime data) that is not already stored in the memory cache 108, then the processing unit 107 initiates a "cache miss" and one or more data blocks are fetched from the system memory 103 and flowed into the memory cache 108, possibly "pushing" it out to the system memory 103 to replace some other data already stored in the memory cache 108.

[0033]

[0039] As shown, persistent storage 104 stores computer-executable instructions and / or data structures representing executable software components. Correspondingly, during execution of this software on processor 102, one or more portions of these computer-executable instructions and / or data structures are loaded into system memory 103. For example, persistent storage 104 is shown as storing computer-executable instructions and / or data structures corresponding to debug component 109, and possibly also computer-executable instructions and / or data structures corresponding to one or more of tracer component 110, emulation component 111, or application(s) 112. In some embodiments, persistent storage 104 also stores data such as replayable execution traces 113 (e.g., generated by tracer component 110 using one or more of the historical debugging techniques described above) and block list(s) 114 generated and / or used by debug component 109.

[0034]

[0040] In some embodiments, under the direction of the debug component 109, the tracer component 110 records or "traces" the execution of the application 112 into one or more replayable execution traces 113. In some embodiments, the tracer component 110 records the execution of the application 112 when it is "live" execution directly on the processor 102, when it is "live" execution on the processor 102 with a managed runtime, and / or when it is emulated execution via the emulation component 111. Thus, FIG. 1 also shows that in some embodiments, the debug component 109 and the tracer component 110 are loaded into the system memory 103 (i.e., the debug component 109' and the tracer component 110'). The arrow between the tracer component 110' and the replayable execution traces 113' indicates that the tracer component 110' records the trace data into one or more replayable execution traces 113', which can then be persistently stored in the persistent storage 104 as one or more replayable execution traces 113.

[0035]

[0041] FIG. 5 illustrates an example of an execution trace. In particular, FIG. 5 illustrates an execution trace 500 that includes multiple data streams (i.e., data streams 501a-501n). In some embodiments, each data stream represents the execution of a different execution context, such as a different thread executed from application 112. In one example, data stream 501a records the execution of a first thread of application 112, and data stream 501n records an nth thread of application 112. As illustrated, data stream 501a includes multiple data packets 502. These data packets are illustrated as having different sizes because the specific data logged in each of the data packets 502 may differ. In some embodiments, when using time travel debugging techniques, one or more of the data packets 502 record inputs (e.g., register values, memory values, etc.) to one or more executable instructions executed as part of this first thread of application 112. In some embodiments, the memory values ​​are obtained as inputs to memory cache 108 and / or as uncached read values. In some embodiments, data stream 501 a also includes one or more key frames (e.g., key frames 503 a and 503 b), each capturing sufficient information (e.g., a snapshot of register and / or memory values) to enable replay of a previous execution of a thread, starting from the time of the key frame and proceeding forward.

[0036]

[0042] In some embodiments, the execution trace also includes the actual code that was executed. Thus, in FIG. 5, each data packet 502 is shown as including a data input portion 504 (white) and a code portion 505 (shaded). In some embodiments, the code portion 505 of each data packet 502 includes executable instructions that are executed based on the corresponding data input, if any. However, in other embodiments, the execution trace omits the actual code that was executed and instead relies on having separate access to the executable code (e.g., a copy of the application 112). In these other embodiments, each data packet specifies an address or offset to the appropriate executable instruction in the application binary image. Although not shown, the execution trace 500 can include a data stream that stores one or more of the outputs of the code execution. It should be noted that the use of different data inputs and code portions of data packets is merely exemplary and the same data can be stored in various ways, such as using multiple data packets.

[0037]

[0043] When there are multiple data streams, in some embodiments, these data streams may include sequence events. Each sequence event records the occurrence of an event that can be ordered across various execution contexts, such as threads. In one example, a sequence event corresponds to an interaction between threads, such as an access to memory shared by the threads. Thus, for example, if a first thread traced to a first data stream (e.g., data stream 501a) writes to a synchronization variable, a first sequence event is recorded in that data stream (e.g., data stream 501a). Then, if a second thread traced to a second data stream (e.g., data stream 501n) reads from that synchronization variable, a second sequence event is recorded in that data stream (e.g., data stream 501n). These sequence events are inherently ordered. For example, in some embodiments, each sequence event is associated with a monotonically increasing value, which defines an overall order between the sequence events. In one example, a first sequence event recorded in the first data stream is given a value of 1, a second sequence event recorded in the second data stream is given a value of 2, and so on.

[0038]

[0044] In some embodiments, under the direction of debug component 109, emulation component 111 emulates the execution of code of an executable entity, such as application 112, based on execution state data obtained from one of replayable execution traces 113. Thus, Figure 1 illustrates that, in some embodiments, debug component 109 and emulation component 111 are loaded into system memory 103 (i.e., debug component 109' and emulation component 111'), and the execution of application 112 is emulated within emulation component 111' (i.e., application 112').

[0039]

[0045] In some embodiments, computer system 101 is part of a networked computing environment, where computer system 101 is connected (e.g., using network device 105) to one or more remote computer systems, each remote computer system including one or more of a corresponding debug component, a corresponding tracer component, or a corresponding emulation component. For example, FIG. 4 illustrates an exemplary computing environment 400, showing computer system 101 of FIG. 1 as connected to remote computer system(s) (e.g., remote computer system 402a-remote computer system 402n) via network(s) 401. In one embodiment, computer system 101 receives one or more replayable execution traces 113 via network 401 from one or more remote computer systems, each of which includes a tracer component (e.g., for analysis of those traces by debug component 109 at computer system 101). In another embodiment, computer system 101 records one or more replayable execution traces 113 using tracer component 110 and transmits those traces over network 401 to one or more remote computer systems that each include one or more debugging or emulation components (e.g., for analysis of those traces by the remote computer systems). In some embodiments, computer system 101 also receives and transmits block lists 114 from one or more remote computer systems over network 401.

[0040]

[0046] It should be noted that in some embodiments, debug component 109, tracer component 110, and / or emulation component 111 are each separate components or applications, while in other embodiments they are alternatively integrated into the same application (e.g., a debug suite) or integrated into another software component (e.g., an operating system component, a hypervisor, a cloud fabric, etc.). Thus, one skilled in the art will appreciate that the present invention may also be implemented in a cloud computing environment that includes computer system 101. For example, in some embodiments, these components take the form of one or more software applications that run on a user's local computer, while in other embodiments they take the form of services provided by a cloud computing environment.

[0041]

[0047] In some embodiments, the debug component 109 is a tool (e.g., a debugger, a profiler, a cloud service, etc.) that consumes one or more replayable execution traces 113 as part of an analysis of a previous execution of the application 112. As described in more detail herein in connection with Figures 2 and 6, in some embodiments, the debug component 109 provides functionality for using entropy analysis to prevent payload data from being included in code execution log data, such as the replayable execution traces 113. Additionally or alternatively, as described in more detail in connection with Figures 3 and 7, in some embodiments, the debug component 109 provides functionality for removing payload data from the replayable execution traces 113. I. Using Entropy to Prevent Payload Data from Being Included in Code Execution Log Data

[0048] It was noted above that some embodiments of the debug component 109 provide functionality for using entropy analysis to prevent inclusion of payload data in code execution log data. These embodiments address the challenge that sensitive data items (e.g., PII, cryptographic keys, passwords, etc.) can be very difficult to programmatically identify by leveraging the observation that sensitive data items tend to have high entropy. Using this observation, these embodiments identify high entropy data items, such as inputs and outputs to executable instructions, functions, etc., and exclude them from inclusion in code execution log data, such as replayable execution traces and event logs.

[0042]

[0049] As used herein, a data item is considered to have "high entropy" when the intrinsic and / or contextual entropy of the data item meets a predetermined threshold. In some embodiments, a data item is considered to have high intrinsic entropy (and thus high entropy) when the ratio of the number of bits of entropy in the data item to the total number of bits in the data item exceeds a predetermined threshold. In one example, cryptographic keys and passwords tend to have high intrinsic entropy because they tend to contain a relatively large number of unique characters relative to the total number of characters in the key / password. Conversely, natural language texts tend to have low intrinsic entropy because they tend to contain a relatively small number of unique characters relative to the total number of characters in the text. In particular, data items with high intrinsic entropy have relatively low internal character repetition when compared to data items with low intrinsic entropy, and therefore data items with high intrinsic entropy have low compression ratios when compared to data items with high intrinsic entropy. Conversely, data items with low intrinsic entropy have relatively high internal character repetition when compared to data items with high intrinsic entropy, and therefore data items with low intrinsic entropy have high compression ratios when compared to data items with low intrinsic entropy. Thus, in some embodiments, the compression ratio of a data item is used to determine whether the data item is considered to have low or high intrinsic entropy.

[0043]

[0050] In some embodiments, a data item is considered to have high contextual entropy (and thus high entropy) when the ratio of the number of times a particular value of the data item occurs to the total number of times the data item appears in a collection of execution traces is below a predetermined threshold, or when the ratio of the number of traces in which a particular value of the data item occurs to the total number of traces in the collection of execution traces is below a predetermined threshold. In one example, social security numbers tend to have high contextual entropy because, given a large collection of data collected across multiple computer systems, a given social security number tends to occur rarely. Conversely, state names tend to have low contextual entropy because, given a large collection of data collected across multiple computer systems, a given state name tends to occur frequently.

[0044]

[0051] In some embodiments, the computer system that generates the code execution log data takes action to exclude high entropy data from being included in the code execution log data. In one example, during generation of one or more replayable execution traces 113 by the tracer component 110 at the computer system 101, the debug component 109 at the computer system 101 uses the block list 114 and / or intrinsic entropy analysis to identify high entropy payload data items and exclude those payload data items from those traces. In another example, during post-processing of one or more replayable execution traces 113 generated by the debug component 109 at the computer system 101, the debug component 109 at the computer system 101 uses the block list 114 and / or intrinsic entropy analysis to identify high entropy payload data items and exclude those payload data items from those traces. In any case, the high entropy payload data items are excluded from the code execution log data before the log data is exported from the computer system 101 (e.g., to a remote computer system 402a).

[0045]

[0052] In other embodiments, the computer system processing the collection of code execution log data takes action to exclude the inclusion of the high entropy data in future collections by identifying code that interacts with the high entropy data in the block list 114 and by sending the block list 114 to one or more remote computer systems. In these embodiments, the computer system may also remove the high entropy data items from the collection of code execution log data used to identify the high entropy data items. In one example, while the debugging component 109 at the computer system 101 processes multiple replayable execution traces 113 of the application 112 received from multiple remote computer systems, the debugging component 109 identifies high entropy payload data items in the multiple execution traces using intrinsic entropy analysis and / or contextual entropy analysis. The debug component 109 then adds those data items to a block list 114 (or a block list 114' in the system memory 103) based on references to code that interacted with those data items, and distributes this block list 114 to remote computer systems to prevent those data items from being included in future execution traces collected from those remote computer systems. In some embodiments, the debug component 109 also removes these high entropy payload data items from the execution traces.

[0046]

[0053] To further demonstrate these concepts, FIG. 2 illustrates an example 200 of a debug component 201 (e.g., one embodiment of debug component 109) configured to use entropy to prevent payload data from being included in code execution log data, including components (e.g., log data interaction component 202, high entropy payload identification component 203, code interaction component 204, block list interaction component 205, preventive action component 206, export component 207, etc.) that operate to use entropy to prevent payload data from being included in code execution log data. The illustrated components of debug component 201, including subcomponents, are representative of various functions that may be implemented or utilized by debug component 201 in accordance with various embodiments described herein. However, it should be understood that the illustrated components (including their identities, subcomponents, and arrangements) are presented only to aid in explaining various embodiments of debug component 201 described herein, and that these components are not limiting as to how software and / or hardware may implement various embodiments of debug component 201, or specific functions thereof, as described herein.

[0047]

[0054] The debug component 201 is described in relation to FIG. 6. FIG. 6 illustrates a flow diagram of a method 600 for using entropy to prevent inclusion of payload data in code execution log data. Accordingly, the following discussion refers to methods and method operations. Although the method operations may be discussed in a particular order or illustrated in a flow diagram as occurring in a particular order, no particular order is required unless otherwise specified or unless an order is required because an operation is dependent on another operation being completed before the operation is performed. In some embodiments, instructions for implementing the method 600 are encoded as computer-executable instructions (e.g., the debug component 201) stored in a hardware storage device (e.g., the persistent storage 104) that are executable by a processor (e.g., the processor 102) to cause a computer system (e.g., the computer system 101) to perform the method 600.

[0048]

[0055] The log data interaction component 202 interacts with code execution log data, e.g., replayable execution traces 113, or other code execution logs, such as event logs. As shown, in various implementations, the log data interaction component 202 includes one or more of a log data access component 202a (i.e., for accessing existing code execution log data, such as replayable execution traces 113 stored in persistent storage 104), a log data generation component 202b (i.e., for generating code execution log data, such as replayable execution traces 113), or a log data modification component 202c (i.e., for modifying existing code execution log data, such as replayable execution traces 113 stored in persistent storage 104). While the type of code execution log data acted upon by the log data interaction component 202 may vary, in some embodiments, the code execution log data is a replayable execution trace, such as one of the replayable execution traces 113.

[0049]

[0056] The high entropy payload identification component 203 identifies high entropy payload data items in connection with generation of code execution log data (e.g., by the tracer component 110) or in connection with post-processing of code execution log data. The code interaction component 204 identifies code that has interacted with the high entropy payload data items by analyzing the interaction of the high entropy payload data items with the executed code or by consulting the block list 114.

[0050]

[0057] 6, method 600 includes operation 601 of identifying a high entropy payload data item associated with the code execution log data, and operation 602 of identifying a particular code that interacted with the high entropy payload data item. Operation 601 and operation 602 are shown without any particular order between the operations. In some embodiments, operation 601 first identifies the high entropy payload data item, and then operation 602 identifies the particular code that interacted with the high entropy payload data item. In other embodiments, the operation identifies the particular code that interacted with the high entropy payload data item from the block list 114, and then operation 601 identifies the high entropy payload data item as the item with which the particular code interacted.

[0051]

[0058] In some embodiments, operation 601 includes determining that a payload data item associated with the code execution log data has entropy above a predefined entropy threshold. In some embodiments, operation 602 includes identifying particular executable code that interacted with the payload data item. In some embodiments, operations 601 and 602 have the technical effect of identifying payload data that may be sensitive data, such as PII, cryptographic keys, passwords, etc., along with code that interacted with the payload data. In particular, identifying payload data that may be sensitive data in operation 601 and identifying code that interacted with the sensitive data in operation 602 may improve data security by allowing this high entropy payload data to be excluded from the code execution log data.

[0052]

[0059] As shown, the high entropy payload identification component 203 includes thresholds 203c that define when the calculated entropy for a given data item is considered to be “high entropy.” As will be appreciated in light of the discussion of intrinsic and contextual entropy, in some embodiments, these thresholds 203c are ratio-based.

[0053]

[0060] As also shown, in some embodiments, the high entropy payload identification component 203 includes an intrinsic entropy component 203a. In some embodiments, the intrinsic entropy component 203a analyzes the payload data based on its intrinsic entropy. Thus, in some embodiments of operation 601, determining that a payload data item has entropy that exceeds a predefined entropy threshold includes determining that the payload data item has intrinsic entropy that exceeds a predefined entropy threshold.

[0054]

[0061] As discussed, a data item has high intrinsic entropy (and therefore high entropy) if the ratio of the number of bits of entropy in the data item to the total number of bits in the data item exceeds a predefined threshold (i.e., threshold 203c). Thus, in some embodiments of operation 601, determining that a payload data item has intrinsic entropy that exceeds a predefined entropy threshold comprises calculating the ratio of the number of bits of entropy in the payload data item to the total number of bits in the payload data item. As also discussed, in some embodiments, the compression ratio of the data item is used to determine whether the data item has low intrinsic entropy (i.e., when it has a relatively high compression ratio) or high intrinsic entropy (i.e., when it has a relatively low compression ratio). Thus, in some embodiments of operation 601, determining that a payload data item has intrinsic entropy that exceeds a predefined entropy threshold comprises calculating the compression ratio of the payload data item.

[0055]

[0062] As also shown, in some embodiments, the high entropy payload identification component 203 includes a context entropy component 203b. In some embodiments, the context entropy component 203b analyzes the payload data based on its context entropy. In some embodiments, a data item has high context entropy (and thus high entropy) when a ratio of the number of times a particular value of the data item occurs to the total number of times the data item appears in the collection of execution traces is less than a predetermined threshold (i.e., threshold 203c). In other embodiments, a data item has high context entropy (and thus high entropy) when a ratio of the number of traces in which a particular value of the data item occurs to the total number of traces in the collection of execution traces is less than a predetermined threshold (i.e., threshold 203c). Thus, in some embodiments of operation 601, determining that a payload data item has entropy that exceeds a predetermined entropy threshold includes determining that the payload data item has context entropy that exceeds a predetermined entropy threshold, the context entropy relating to a plurality of associated payload data items identified from a plurality of associated code execution logs.

[0056]

[0063] In some embodiments, when identifying a high entropy payload data item in operation 601, the high entropy payload identification component 203 utilizes the block list interaction component 205 to identify known high entropy payload data from a block list 114, such as a block list received from a remote computer system over the network 401. As discussed, in some embodiments, the block list 114 identifies high entropy data items with reference to code that interacted with the data item, such as code that consumed or generated the data item. Thus, in some embodiments of operation 601, determining that a payload data item has entropy above a predefined entropy threshold includes identifying the payload data item from the block list based at least on a reference to particular executable code (identified in operation 602) that interacted with the payload data item.

[0057]

[0064] In some embodiments, the high entropy payload identification component 203 operates during post-processing of the code execution log data, e.g., by a computer system that generated the code execution log data or by a computer system that receives the code execution log data from another computer system. Thus, in some embodiments of operation 601, determining that a payload data item has entropy above a predefined entropy threshold includes identifying the payload data item during post-processing of the code execution log data. In other embodiments, the high entropy payload identification component 203 operates during generation of the code execution log data, such as during trace recording by the tracer component 110. Thus, in some embodiments of operation 601, determining that a payload data item has entropy above a predefined entropy threshold includes identifying the payload data item during generation of the code execution log data.

[0058]

[0065] As discussed above, in operation 602, the code interaction component 204 identifies particular executable code that interacted with the payload data item. In various embodiments, this interaction may be consumption of the payload data item (wherein identifying the particular executable code that interacted with the payload data item in operation 602 includes identifying the executable code that consumed the payload data item) or production of the payload data item (wherein identifying the particular executable code that interacted with the payload data item in operation 602 includes identifying the executable code that generated the payload data item). The granularity at which the code interaction component 204 identifies the executable code may vary, such as at the instruction level (wherein identifying the particular executable code that interacted with the payload data item in operation 602 includes identifying a particular executable instruction) or at the function level (wherein identifying the particular executable code that interacted with the payload data item in operation 602 includes identifying a particular function), etc.

[0059]

[0066] The preventive action component 206 takes preventive action to prevent export of payload data (e.g., by the export component 207), to prevent the payload data item from being included in code execution log data, and / or to add the payload data item to a block list. Turning to Figure 6, the method 600 includes an operation 603 that takes an action to exclude the high entropy payload data item from records of execution of the particular code. In some embodiments, operation 603 includes taking preventive action to exclude the payload data item from being included in records of execution of the particular executable code.

[0060]

[0067] As discussed above, in some embodiments, the preventive action component 206 takes preventive action to prevent the export of payload data. Thus, as shown in FIG. 6, in some embodiments, operation 603 includes operation 603a that prevents the export of a payload data item, and the preventive action at operation 603 includes preventing the payload data item from being exported from the computer system. For example, in some embodiments, the preventive action component 206 removes high entropy payload data items (identified by the high entropy payload identification component 203 at operation 601) from the code execution log data when the code execution log data is exported to a remote computer system by the export component 207.

[0061]

[0068] As discussed above, in some embodiments, the preventive action component 206 takes preventive action to prevent payload data items from being included in the code execution log data. Thus, as shown in FIG. 6, in additional or alternative embodiments, operation 603 includes operation 603b that prevents the payload data item from being included in the code execution log data, and the preventive action at operation 603 includes preventing the payload data item from being included in the code execution log data. For example, in some embodiments, the preventive action component 206 prevents high entropy data items (identified by the high entropy payload identification component 203 in operation 601) from being included in currently generated code execution log data and / or removes those high entropy data items from existing code execution log data.

[0062]

[0069] In various embodiments, preventing the payload data item from being included in the code execution log data includes replacing the payload data item in the code execution log data with one or more of: (i) alternative data, (ii) one or more constraints on the payload data item, or (iii) a code flow override for the particular executable code. These techniques are discussed in more detail in connection with Figures 3 and 7, which describe methods for achieving removal of payload data from an execution trace.

[0063]

[0070] As discussed above, in some embodiments, the preventive action component 206 performs a preventive action of adding payload data items to a block list. Thus, as shown in FIG. 6, in additional or alternative embodiments, operation 603 includes operation 603c of adding payload data items to a block list, where the preventive action in operation 603 includes adding the payload data items to the block list with reference to specific executable code, where the block list is structured to prevent the payload data items from being included in subsequently generated code execution log data. For example, in some embodiments, the preventive action component 206 uses the block list interaction component 205 to add high entropy data items (identified by the high entropy payload identification component 203 in operation 601) to the block list 114 with reference to code (identified by the code interaction component 204 in operation 602) that interacts with those data items. Thus, operation 603a has the effect of preventing the high entropy data items from being included in code execution log data that has not yet been generated.

[0064]

[0071] In some embodiments, the debug component 201 operates transiently when preventing inclusion of payload data in the code execution log data, thus excluding multiple instances of high entropy payload data (including derivatives) from being included in the code execution log data. For example, in some embodiments, when the high entropy payload identification component 203 determines that the data of a first parameter of a particular function is a high entropy value, it also identifies any additional code locations that also interact with that data (or copies / derivatives thereof). The preventive action component 206 then takes preventive action with respect to each of these identified code locations to exclude this data from being included in the code execution log data. For example, if the first parameter (or a derivative thereof) is later printed to the screen as a string, in some embodiments the high entropy payload identification component 203 identifies code that prints this string payload and that clears the string payload. In some embodiments, this is true even if there is no print routine in the recorded code path. Thus, in some embodiments, the debug component 201 operates to transiently exclude all instances of the high entropy payload data (including derivatives thereof) from the code execution log data. Thus, in some embodiments of the method 600, the method 600 is applied transiently to exclude additional instances of the payload data item or derivatives thereof from being included in another execution record of another executable code.

[0065]

[0072] In some embodiments, even though payload data items are prevented from being exported by export component 207, prevented from being included in the code execution log data, and / or removed from the code execution log data, in some embodiments preventive action component 206 separately retains these payload data items (e.g., in persistent storage 104) and possibly encrypts or otherwise protects their data times. In some embodiments, this retention enables their data times to be provided at a later time, should the need arise, to facilitate analysis of the code execution log data. Thus, in some embodiments of method 600, the payload data items are retained in a computer system separately from the code execution log data.

[0066]

[0073] In particular, high entropy data may be excluded from the code execution log data to improve data security, for example by excluding the high entropy data from the code execution record by preventing the export of the payload data item in operation 603a, by preventing the inclusion of the payload data item in the code execution log data in operation 603b, or by adding the payload data item to a block list in operation 603c. This has the technical effect of preventing the inadvertent exposure of sensitive data items via the code execution log data. Furthermore, excluding the high entropy payload data items from the code execution record reduces the size of the code execution log data by removing portions of the data, including removing high entropy data that is often not easily compressible, such as the replayable execution traces 113, which has the additional technical effect of saving computing resources. For example, reducing the size of the replayable execution traces 113 saves processing resources when analyzing those traces and saves storage and network resources required to store and transfer those traces. II. Removing payload data from execution traces

[0074] It was noted above that some embodiments of the debug component 109 additionally or alternatively provide functionality for removing payload data from the replayable execution trace 113. These embodiments are based on the recognition that payload data frequently does not contribute substantially to many forms of execution trace analysis, such as code and / or data flow analysis, memory locality analysis, memory access pattern analysis, cache usage analysis, race condition analysis, buffer overflow analysis, traditional debugging, etc. By removing this payload data from the execution trace, these embodiments improve data security (i.e., by removing such data execution trace) and also significantly reduce trace file size.

[0067]

[0075] Conceptually, these embodiments transform an execution trace from one that explicitly or inherently captures all inputs and all outputs to each instruction executed (i.e., payload data) to one that includes the actual code that was executed, information about what caused that code to be executed as it was, and memory access patterns (i.e., memory addresses accessed). In particular, these embodiments process the execution trace to identify payload data items, such as inputs to and outputs from executable instructions, functions, etc. For each identified payload data item, these embodiments determine constraints that execution of code that interacted with the payload data item imposed on the payload data item. These embodiments then replace the value of the payload data item in the execution trace with information that maintains those constraints. In some embodiments, this information includes the actual executable code that interacted with the payload data item, the memory addresses of the payload data item, and / or data structured to preserve the code flow affected by the payload data item (e.g., replacement values ​​for payload data items that preserve the code flow, specifications of valid values ​​for data items that preserve the code flow, instructions of the code path to follow, etc.).

[0068]

[0076] To further demonstrate these concepts, FIG. 3 illustrates an example 300 of a debug component 301 (e.g., an embodiment of the debug component 109) configured to remove payload data from an execution trace, including components that operate to remove payload data from an execution trace (e.g., a trace interaction component 302, a payload identification component 303, a code interaction component 304, a constraint identification component 305, a payload replacement component 306, etc.). The illustrated components of the debug component 301, including subcomponents, are representative of various functions that the debug component 301 may implement or utilize in accordance with various embodiments described herein. However, it should be understood that the illustrated components (including their identities, subcomponents, and arrangements) are presented only to aid in explaining various embodiments of the debug component 301 described herein, and that these components are not limiting as to how software and / or hardware may implement various embodiments of the debug component 301, or specific functions thereof, as described herein.

[0069]

[0077] The debug component 301 is described in relation to FIG. 7. FIG. 7 illustrates a flow diagram of a method 700 for removing payload data from an execution trace. Accordingly, in the following discussion, reference is made to methods and method operations. Although the method operations may be discussed in a particular order or may be illustrated in a flow diagram as occurring in a particular order, no particular order is required unless otherwise specified or unless an order is required because an operation is dependent on another operation being completed before the operation is performed. In some embodiments, instructions for implementing the method 700 are encoded as computer-executable instructions (e.g., debug component 301) stored in a hardware storage device (e.g., persistent storage 104) that are executable by a processor (e.g., processor 102) to cause a computer system (e.g., computer system 101) to perform the method 700.

[0070]

[0078] The trace interaction component 302 interacts with execution traces, such as the replayable execution traces 113. As shown, the trace interaction component 302 includes a trace access component 302a (i.e., for accessing existing execution traces) and a trace modification component 302b (i.e., for modifying existing execution traces).

[0071]

[0079] The payload identification component 303 identifies payload items from the execution trace accessed by the trace access component 302a. In some embodiments, the payload identification component 303 identifies inputs to and outputs from code instructions, functions, modules, etc. In some embodiments, the payload identification component 303 also identifies a memory address (or even a range of memory addresses) corresponding to each payload item. In some embodiments, the payload identification component 303 also identifies a name or label for each payload item, such as a variable name, a structure name, a class name, etc. Referring to FIG. 7, the method 700 includes an operation 701 of identifying a payload data item from the execution trace. In some embodiments, the operation 701 includes identifying a payload data item that is in one of the replayable execution traces 113 accessed by the trace interaction component 302.

[0072]

[0080] The code interaction component 304 identifies code that interacted with each payload data item identified by the payload identification component 303. As previously described, the payload identification component 303 identifies inputs to and outputs from code instructions, functions, modules, etc. Thus, the code interaction component 304 identifies the code of these instructions, functions, modules, etc. In some embodiments, the code interaction component 304 further identifies from which memory address this code was accessed. Turning to FIG. 7, the method 700 includes an operation 702 of identifying a particular code that interacted with the payload data item. In some embodiments, the operation 702 includes identifying a particular executable code that interacted with the payload data item identified by the payload identification component 303 in operation 701.

[0073]

[0081] In the example of operations 701 and 702, the execution trace records the execution of a string copy function as follows:

[0074]

number

[0075] [Table 1] Now, after execution of the strcpy() function, these same characters are copied to another block of memory pointed to by *dest. As an example, when processing this execution trace, for each iteration of the while loop, the payload identification component 303 identifies one of the characters of the "Sample" string as an input to a branch instruction (e.g., corresponding to while), an input to a memory load instruction, and an input to a memory store instruction, along with the corresponding memory addresses for those payload items. Correspondingly, the code interaction component 304 identifies these instructions as code that interacted with these payload data items, along with the corresponding memory addresses for that code.

[0076]

[0082] The constraint identification component 305 identifies constraints that execution of the code identified by the code interaction component 304 imposes on the payload identified by the payload identification component 303. Turning to Figure 7, the method 700 includes an operation 703 of determining constraints that execution of the particular code imposes on the payload data item. In some embodiments, the operation 703 includes determining, based on the payload data item and the particular executable code, one or more constraints that execution of the particular executable code imposes on the payload data item.

[0077]

[0083] As shown, in some embodiments, the constraint identification component 305 comprises a code bytes component 305a. As will be appreciated, the actual code that interacts with the payload data is an inherent constraint on that payload data as a code input. Thus, in some embodiments, the code bytes component 305a identifies the actual bytes of the executed code for inclusion in the execution trace. In some embodiments, the code bytes component 305a identifies these code bytes based on data stored in the code memory addresses identified by the code interaction component 304. As shown in FIG. 7, in some embodiments, the operation 703 includes an operation 703a that identifies code bytes. Thus, in some embodiments of the operation 703, the one or more constraints include one or more bytes of a particular executable code. In one example, when operating on the strcpy() example above, the code bytes component 305a identifies the actual code of the strcpy() function.

[0078]

[0084] As also shown, in some embodiments, the constraint identification component 305 comprises a memory address component 305b. As will be appreciated, the location where the executable code accessed the payload data imposes location constraints on the payload data. Thus, in some embodiments, the memory address component 305b identifies memory addresses at which payload data items were accessed for inclusion in the execution trace. As shown in FIG. 7, in some embodiments, the operation 703 includes an operation 703b that identifies memory addresses. Thus, in some embodiments of the operation 703, one or more constraints include memory addresses corresponding to the payload data items. In one example, when operating on the strcpy() example above, the memory address component 305b identifies memory addresses of each character in the src and dest strings.

[0079]

[0085] As also shown, in some embodiments, the constraint identification component 305 comprises a code flow component 305c. As will be appreciated, code flows resulting from interactions with payload data impose constraints on possible values ​​of the payload data that preserve the same code flow. Thus, in some embodiments, the code flow component 305c identifies code flow constraints on the payload data for inclusion in the execution trace. As shown in FIG. 7, in some embodiments, operation 703 includes operation 703c that identifies data structured to preserve code flow. Thus, in some embodiments of operation 703, the one or more constraints include data structured to preserve code flow.

[0080]

[0086] In some embodiments, the code flow component 305c identifies data structured to preserve the code flow in the form of a replacement value for the payload data item, the replacement value being a value that preserves the appropriate code flow when replacing the original value of the payload data item. Thus, in some embodiments, operation 703c includes a replacement value for the payload data item, the replacement value being structured to preserve the code flow.

[0081]

[0087] In some embodiments, the code flow component 305c selects this replacement value based on a random value generation, and operation 703c includes identifying the replacement value based on the random value generation. In some embodiments, when the code flow component 305c selects the replacement value based on a random value, the code flow component 305c generates a random value of an appropriate data size for the payload data item (e.g., using a pseudo-random value generation technique) and then checks whether the code flow would be preserved if that randomly generated value was used as the value of the payload data item. If the code flow is preserved, the code flow component 305c selects the randomly generated value as the replacement value. If the code flow is not preserved, the code flow component 305c generates and checks new random values ​​until a value that preserves the code flow is identified. In particular, in some embodiments, the code flow component 305c enables the generation of a random value identical to the original value of the payload data item. In one example, when operating on the strcpy() example above, the code flow component 305c generates a different random one-byte value for each alphabetic character in the "Sample" string. Here, non-null bytes (i.e. bytes other than 0x0) preserve the code flow.

[0082]

[0088] In additional or alternative embodiments, the code flow component 305c selects the replacement value based on a search from a set of available replacement values, and operation 703c includes identifying the replacement value based on a search from the set of available replacement values. In some embodiments, when the code flow component 305c selects a replacement value based on a search, the code flow component 305c selects a value from a predefined set of available replacement values ​​and then checks whether that value would preserve the code flow if used as the value of the payload data item. If the code flow would be preserved, the code flow component 305c selects the selected value as the replacement value. If the code flow would not be preserved, the code flow component 305c selects and checks other values ​​from the set until a value that preserves the code flow is identified. In particular, in some embodiments, the code flow component 305c enables the selection of a value that is identical to the original value of the payload data item. In some embodiments, the set of available replacement values ​​is populated with values ​​(e.g., 0x0, 0xA, 0xF, etc.) that promote compressibility of execution traces that include many instances of those values. In one example, when operating on the strcpy() example above, the code flow component 305c generates a possible one-byte value for each character in the string from a set of possible values, where any non-null bytes such as 0xA and 0xF preserve the code flow.

[0083]

[0089] In additional or alternative embodiments, the code flow component 305c selects this replacement value based on at least a calculation of a hash of the payload data item, and operation 703c includes generating the replacement value based on a calculation of a hash of the payload data item. In some embodiments, the code flow component 305c checks whether the generated hash would preserve the code flow if used as a replacement value for the payload data item, and if so, the code flow component 305c selects the hash as the replacement value. Thus, in some embodiments of operation 703c, replacing the value of the payload data item in the execution trace with information that maintains one or more constraints includes replacing the value of the payload data item with the hash. In some embodiments, if the generated hash does not preserve the code flow if used as a replacement value for the payload data item, the code flow component 305c uses one of the other replacement value selection techniques discussed herein (e.g., a randomly generated value, a value selected from a set of available replacement values, etc.) and tags the replacement value with the hash. Thus, in some embodiments of operation 703c, replacing the value of the payload data item in the execution trace with information that maintains the one or more constraints includes tagging the replacement value with a hash. In one example, when operating on the strcpy() example above, the code flow component 305c generates a hash of each letter in the string, where these hashes are non-null and therefore preserve the code flow and are used as replacements for each letter in the string.

[0084]

[0090] In some embodiments, when the code flow component 305c selects a replacement value based on a hash, the code flow component 305c hashes the value of the payload data item, either alone or in combination with a salt. Thus, in some embodiments of operation 703c, computing the hash of the payload data item includes applying a salt to the payload data item. In some embodiments, the code flow component 305c uses the same salt when hashing values ​​within a set of traces being analyzed for a common purpose, such as analyzing a particular bug or flaw. In this way, payloads having the same value are replaced (or at least tagged) with the same hash across this set of traces and can therefore be associated with each other during analysis. However, if a different salt is used for another set of traces, the overall obfuscation of the payload across all traces is preserved, since they will have different hashes when different salts are used.

[0085]

[0091] In additional or alternative embodiments, the code flow component 305c selects this replacement value based on execution of a constraint solver (e.g., using Boolean Satisfiability Problem (SAT) techniques), and operation 703c includes identifying the replacement value based on execution of the constraint solver for at least the particular executable code. In some embodiments, when the code flow component 305c selects the replacement value based on a constraint solver, the code flow component 305c uses the constraint solver to analyze the code that interacted with the payload data item. In one example, when operating on the strcpy() example above, the code flow component 305c runs the constraint solver on the code that makes up the while loop.

[0086]

[0092] In some embodiments, the code flow component 305c identifies data structured to store a code flow in the form of a specification of a set of one or more valid values ​​for a payload data item. Thus, in some embodiments of operation 703c, the data structured to store a code flow includes a specification of a set of one or more valid values ​​for a payload data item. The specification of the set of one or more valid values ​​can take various forms, such as a range of values ​​(e.g., 1-10), a set of values ​​(e.g., 1, 2, 4, 8, and 10), a flag (e.g., zero or non-zero), a boundary value (e.g., less than 7), etc. In one example, when operating on the strcpy() example above, the code flow component 305c specifies that each alphabetic character in the "Sample" string can be replaced with a non-null character.

[0087]

[0093] In some embodiments, the code flow component 305c identifies data structured to store a code flow in the form of instructions of a code path to be followed. Thus, in some embodiments of operation 703c, the data structured to store a code flow includes instructions of a code path to be followed within a particular executable code. The instructions of a code path to be followed may include, for example, instructions on which branch of a control statement should be taken, instructions on whether a statement should evaluate to true or false, etc. In one example, when operating on the strcpy() example above, for each alphabetic character in the "Sample" string, the code flow component 305c indicates that the while loop should continue, and for a null termination character, indicates that the while loop should end. As will be appreciated, by specifying instructions of a code path to be followed, the associated payload data may be omitted entirely from the trace.

[0088]

[0094] The payload replacement component 306 performs payload data replacement based on the constraints identified by the constraint identification component 305 in operation 703. With reference to Figure 7, the method 700 includes an operation 704 of replacing a payload data item with information that maintains the constraints. In some embodiments, the operation 704 includes replacing values ​​of the payload data items in the execution trace with information that maintains the one or more constraints.

[0089]

[0095] As discussed, in some embodiments, the one or more constraints include the actual code (code bytes component 305a) that interacts with the payload data. Correspondingly, in some embodiments, the payload replacement component 306 comprises a code bytes component 306a that logs at least the code bytes that interacted with the payload data to the execution trace. Thus, in some embodiments of operation 704, the information that maintains the one or more constraints includes one or more bytes of a particular executable code. Thus, as shown, in some embodiments, operation 704 includes operation 704a that logs code bytes. In some embodiments, operation 704a has the technical effect of enabling access to the executable code to facilitate execution trace analysis.

[0090]

[0096] As discussed, in some embodiments, the one or more constraints include actual code (code bytes component 305a) that interacts with the payload data. Correspondingly, in some embodiments, the payload replacement component 306 comprises a memory address component 306b that logs at least a memory address corresponding to the payload data. Thus, in some embodiments of operation 704, the information that maintains the one or more constraints includes a memory address. Thus, as shown, in some embodiments, operation 704 includes operation 704b that logs a memory address. In some embodiments, operation 704b has the technical effect of enabling analysis of memory usage (e.g., code and / or data flow analysis, memory locality analysis, memory access pattern analysis, cache usage analysis, race condition analysis, buffer overflow analysis, traditional debugging, etc.) even when the original payload data is not present.

[0091]

[0097] As discussed, in some embodiments, the one or more constraints include a code flow resulting from an interaction with the payload data (code flow component 305c). Correspondingly, in some embodiments, the payload replacement component 306 comprises a code flow component 306c that logs data structured to preserve at least the code flow. Thus, in some embodiments of operation 704, the information that maintains the one or more constraints includes data structured to preserve the code flow, where the code flow may include, by way of example, a replacement value for a payload data item, or instructions for a code path to follow within a particular executable code. Thus, as illustrated, in some embodiments, operation 704 includes operation 704c that logs data structured to preserve the code flow. In some embodiments, operation 704c has the technical effect of replacing potentially sensitive data in the execution trace with data structured to preserve the code flow, which promotes data security. Furthermore, when the replacement data is smaller than the original payload data and / or when the replacement data has a compressible data pattern, operation 704 has the technical effect of reducing the trace file size. Thus, the efficiency of trace analysis, storage and transmission (in terms of required processing power, computation time, storage resources and network utilization) is improved, since there is less data to process, store and transmit.

[0092]

[0098] In some embodiments, similar to debug component 201, debug component 301 operates transiently when removing / replacing payload data from replayable execution trace 113, such that multiple instances of that payload data (including derivatives) are removed / replacing from replayable execution trace 113. For example, suppose the strcpy() function described above is a ToLower() function that includes:

[0093]

number

[0094]

[0099] Thus, at least some embodiments herein operate to remove payload data from execution traces, which has the technical effect of promoting data security by preventing inadvertent exposure of sensitive data items via execution traces. Furthermore, excluding payload data from execution traces can significantly increase the compression ratio of those execution traces. In any case, excluding payload data from execution traces can significantly reduce the size of the execution traces, which has the additional technical effect of saving computing resources. For example, reducing the size of the execution traces saves processing resources when analyzing those execution traces, as well as saving storage and network resources required when storing and transferring the execution traces.

[0095]

[0100] In this disclosure, we have discussed embodiments using entropy analysis to prevent inclusion of payload data in code execution log data, and embodiments for removing payload data from replayable execution traces. It should be noted that these two embodiments can be implemented alone or in combination. For example, when discussing the entropy analysis embodiment, it was stated that preventing a payload data item from being included in the code execution log data can include replacing the payload data item with one or more of alternative data, constraints on the payload data item, or code flow overrides. Notably, these prevention techniques are further disclosed in connection with the payload data removal embodiment. Additionally, the payload data removal embodiment identifies payload data items to be removed, and in some implementations, the payload data items are identified at least in part based on entropy analysis.

[0096]

[0101] Although the subject matter has been described in language specific to structural features and / or methodological operations, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to those features or operations or to the sequence of operations described above. Rather, the above features and operations are disclosed as exemplary forms of implementing the claims.

[0097]

[0102] The present invention may be embodied in other specific forms without departing from its essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. The scope of the present invention is therefore indicated by the appended claims, rather than by the foregoing description. All changes that come within the meaning and range of equivalence of the claims are intended to be embraced within their scope. When introducing elements in the appended claims, the articles "a", "an", "the", and "said" are intended to mean that there are one or more of the elements. The terms "comprise", "include", and "have" are intended to be inclusive and mean that there may be additional elements other than the listed elements. Unless otherwise specified, the terms "set", "superset", and "subset" are intended to exclude empty sets, and thus a "set" is defined as a non-empty set, a "superset" is defined as a non-empty superset, and a "subset" is defined as a non-empty subset. Unless otherwise specified, a "subset" excludes its entire superset (i.e., a superset includes at least one item that is not included in the subset). Unless otherwise specified, a "superset" may include at least one additional element, and a "subset" may exclude at least one element.

Claims

1. 1. A method, implemented in a computer system including a processor, for preventing inclusion of payload data in code execution log data, comprising: determining that a payload data item associated with the code execution log data has entropy that exceeds a predefined entropy threshold; identifying the particular executable code that interacted with the payload data item; taking preventative action to exclude said payload data item from being included in the execution trace of said particular executable code; The method includes:

2. 2. The method of claim 1, wherein determining that the payload data item has entropy that exceeds the predetermined entropy threshold comprises determining that the payload data item has intrinsic entropy that exceeds the predetermined entropy threshold.

3. 3. The method of claim 2, wherein the step of determining that the payload data items have intrinsic entropy exceeding the predetermined entropy threshold comprises: calculating a ratio of the number of bits of entropy in said payload data item to the total number of bits in said payload data item; calculating the compressibility of said payload data items; The method of claim 1, further comprising at least one of:

4. 2. The method of claim 1, wherein determining that the payload data item has entropy above the predetermined entropy threshold comprises determining that the payload data item has contextual entropy above the predetermined entropy threshold, the contextual entropy relating to a plurality of associated payload data items identified from a plurality of associated code execution logs.

5. 2. The method of claim 1, wherein identifying the particular executable code that interacted with the payload data item comprises identifying executable code that consumed the payload data item.

6. 2. The method of claim 1, wherein identifying the particular executable code that interacted with the payload data item comprises identifying the executable code that generated the payload data item.

7. 2. The method of claim 1, wherein identifying the particular executable code that interacted with the payload data item comprises identifying a particular executable instruction.

8. 2. The method of claim 1, wherein identifying the particular executable code that interacted with the payload data item comprises identifying a particular function.

9. The method of claim 1 , wherein the preventive action comprises preventing the payload data item from being exported from the computer system.

10. The method of claim 1 , wherein the preventative action comprises preventing the payload data item from being included in the code execution log data.

11. 11. The method of claim 10, wherein preventing the payload data item from being included in the code execution log data comprises replacing the payload data item in the code execution log data with one or more of: (i) alternative data; (ii) one or more constraints on the payload data item; and (iii) a code flow override for the particular executable code.

12. 2. The method of claim 1, wherein the preventive action comprises adding the payload data item to a block list with reference to the particular executable code, the block list being structured to prevent the payload data item from being included in subsequently generated code execution log data.

13. 2. The method of claim 1, wherein determining that the payload data item has entropy above the predetermined entropy threshold comprises identifying the payload data item from a block list based at least on a reference to the particular executable code.

14. 2. The method of claim 1, wherein determining that the payload data items have entropy above the predetermined entropy threshold comprises identifying the payload data items during post-processing of the code execution log data.

15. 10. The method of claim 1, wherein the method is applied transitively to exclude additional instances of a payload data item or derivatives thereof from being included in another execution record of another executable code.