Protecting sensitive information in time travel tracking commissioning
By identifying and deleting sensitive information in the tracking of the time travel debugger, the problem that the debugger may leak sensitive data is solved, and safe and efficient time travel tracking in a production environment is achieved.
Patent Information
- Application Number
- CN202510130123.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-03-15
- Filing Date
- 2019-03-08
- Publication Date
- 2025-05-30
AI Technical Summary
The time travel debugger may capture and leak sensitive code and data while recording precise bit tracking performed by a program, resulting in security risks.
Ensure that sensitive information is not leaked by identifying and deleting or obscuring it in the tracking, for example by storing alternative data, replacing sensitive code or data, and overwriting the execution of instructions.
It realizes the generation and time-consuming travel tracking in the production environment while keeping sensitive information not leaked, enhancing the security of the system.
Smart Images

Figure CN120066942A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application is a divisional application of a patent application for invention, with an international filing date of March 8, 2019, entering the Chinese national phase on September 14, 2020, with a Chinese national application number of 201980018994.0, and an invention title of "Protecting Sensitive Information in Time Travel Trace Debugging". Background Art
[0003] When writing code during the development of a software application, developers typically spend a significant amount of time "debugging" the code to find runtime errors and other source code errors. In doing so, developers can employ several methods to reproduce and locate source code errors, such as observing the behavior of the program based on different inputs, inserting debug code (e.g., to print variable values, trace execution branches, etc.), temporarily removing sections of code, etc. Tracking runtime errors to identify code errors can consume a significant portion of the application development time.
[0004] To assist developers in the code debugging process, many types of debugging applications ("debuggers") have been developed. These tools enable developers to track, visualize, and change the execution of computer code. For example, a debugger can visualize the execution of code instructions, present the values of code variables at different times during code execution, enable developers to change the code execution path, and / or enable developers to set "breakpoints" and / or "watchpoints" on code elements of interest (which cause the execution of the code to be suspended when these points are reached during execution), etc.
[0005] An emerging form of debugging application implements "time travel", "reverse", or "historical" debugging. With "time travel" debugging, the execution of a program (e.g., an executable entity such as a thread) is recorded / tracked by a tracing application into one or more trace data streams. These trace data streams can then be used to replay the execution of the program at a later time for both forward and backward analysis. For example, a "time travel" debugger can enable developers to set forward breakpoints / watchpoints (like a conventional debugger) as well as reverse breakpoints / watchpoints.
[0006] Because time-travel debuggers record bit-exact traces of program execution (including the code executed and the memory values read during program execution), they have the potential to capture and leak sensitive code and / or data that, in many cases, should not be available to those with access to the resulting trace data (e.g., developers using a debugger that consumes the trace data stream). This can be due to security contexts (e.g., kernel vs. user mode), changes in code authorship (e.g., code developed by one author vs. a call library developed by another author), organizational departments, policy / legal issues, etc. For example, a time-travel debugger can capture password information such as the values of encryption keys, random numbers, salts, hashes, nonces, etc.; personally identifiable information (PII) such as names, mailing addresses, birthdays, social security numbers, email addresses, IP addresses, MAC addresses, etc.; financial information such as credit card numbers, account numbers, financial institutions; authentication information such as usernames, passwords, biometric data, etc.; general inputs such as search terms, file names, etc.; code that may be desired to be kept confidential; and so on. The ability of time-travel debuggers to leak sensitive information is becoming an increasing concern as time-travel debugging techniques are evolving to the point where they can have a low enough recording overhead such that they can be used in production systems and even potentially in an "always-on" configuration. Summary of the Invention
[0007] At least some embodiments described herein identify sensitive information associated with time-travel tracing (during the trace recording and / or at some later time) and remove and / or mask the sensitive information in the trace. For example, embodiments can include storing alternative data in the trace (instead of the original data identified as sensitive), replacing original instructions in the trace with alternative instructions that avoid executing sensitive code or that cause correct execution given the data replacement, overriding the execution behavior of one or more instructions, and so on. In this way, embodiments support the generation and consumption of time-travel traces even in a production environment while keeping sensitive information from being leaked.
[0008] Embodiments may include methods, systems, and computer program products for protecting sensitive information related to the original execution of a traced entity. These embodiments may include, for example, identifying that original information includes sensitive information that is accessed based on the original execution of one or more original executable instructions of the entity. Based on the original information including sensitive information, these embodiments may include performing one or both of the following: (i) storing first trace data including alternative information rather than the original information into a first trace data stream while ensuring that the execution path adopted by the entity based on the original information will also be adopted during replay of the original execution of the entity using the first trace data stream; or (ii) storing second trace data into a second trace data stream that causes one or more alternative executable instructions rather than one or more original executable instructions of the entity to be executed during replay of the original execution of the entity using the second trace data stream.
[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This "Summary" is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To describe the manner in which the above and other advantages and features of the invention can be obtained, the invention briefly described above will be described in more detail with reference to specific embodiments of the invention shown in the drawings. After understanding that these drawings only show typical embodiments of the invention and should not be considered as limiting the scope of the invention, the invention will be described and explained with additional features and details by using the drawings, in which:
[0011] Figure 1 An example computer architecture that facilitates protecting sensitive information related to the original execution of a traced entity is illustrated;
[0012] Figure 2 An example embodiment of a security component is illustrated;
[0013] Figure 3A An example of relying on bit-exact tracing to identify derived data and / or code in the forward direction on an execution timeline is illustrated;
[0014] Figure 3B An example of relying on bit-exact tracing to identify derived data and / or code in the reverse direction on an execution timeline is illustrated;
[0015] Figure 4A An example of sensitive data item replacement / masking with respect to a single trace data stream is illustrated;
[0016] Figure 4B Illustrates an example of sensitive data item replacement / masking for multiple traced data streams;
[0017] Figure 5A Illustrates an example that ensures that even if data replacement occurs, the execution path taken by an entity during its original execution will be taken during replay;
[0018] Figure 5B Illustrates an example of storing data into a single traced data stream that causes alternative executable instructions to be executed during replay of an entity;
[0019] Figure 5C Illustrates an example of storing data into at least one traced data stream that causes alternative executable instructions to be executed during replay of an entity; and
[0020] Figure 6 Illustrates a flowchart of an example method for protecting sensitive information related to tracing the original execution of an entity. DETAILED DESCRIPTION
[0021] At least some embodiments described herein identify sensitive information related to time travel tracing (during trace recording and / or at some later time), and delete and / or mask the sensitive information in the trace. For example, embodiments may include storing alternative data in the trace (instead of the original data identified as sensitive), replacing original instructions in the trace with alternative instructions that avoid executing sensitive code or that cause correct execution considering data replacement, overriding the execution behavior of one or more instructions, and so on. In this way, embodiments support time travel tracing being generated and consumed even in a production environment while keeping sensitive information from being leaked.
[0022] As used in this specification and the claims, phrases such as "sensitive information", "sensitive data", "sensitive code", etc. refer to data and / or code that is consumed in one or more processing units during the tracking of these processing units into one or more tracking data streams and that should (or may should) be restricted and / or prevented from being available to the consumers of these tracking data streams. As described in the background art, sensitive data may correspond to, for example: password information, such as values of encryption keys, random numbers, salts, hashes, nonces, etc.; personally identifiable information (PII), such as name, mailing address, birthday, social security number, email address, IP address, MAC address, etc.; financial information, such as credit card numbers, account numbers, financial institutions; authentication information, such as usernames, passwords, biometric data, etc.; general inputs, such as search terms, file names, etc.; and so on. Sensitive code may correspond to code that executes encryption routines, code that implements proprietary algorithms, etc. The classification of sensitive data or code may be based on security context (e.g., kernel vs. user mode), changes in code authors (e.g., code developed by one author vs. a call library developed by another author), organizational departments, policies, and / or legal issues, etc.
[0023] As used herein, phrases such as "non-sensitive information", "non-sensitive data", "non-sensitive code", etc. refer to information that may be non-sensitive. For example, this may include information for which the confidence that the information is non-sensitive is substantially 0% or below a predetermined threshold (e.g., such as 10%). In contrast, "sensitive information", "sensitive data", "sensitive code", etc. encompass information that is definitely sensitive, possibly sensitive, and potentially sensitive. Thus, unless otherwise stated, the use of the phrases "sensitive information", "sensitive data", "sensitive code" (etc.) should be interpreted broadly to include information that is definitely sensitive, possibly sensitive, and potentially sensitive. In some embodiments, definitely sensitive information may include information for which the confidence that the information is sensitive is substantially 100% or above a predetermined threshold (e.g., such as 95%). In some embodiments, possibly sensitive information may include information for which the confidence that the information is sensitive exceeds a predetermined threshold (e.g., such as >50% or >=75%). In some embodiments, potentially sensitive information may include information for which the confidence that the information is sensitive is between the threshold for non-sensitive information and the threshold for possibly sensitive information.
[0024] Figure 1FIG. illustrates an example computing environment 100 that facilitates protecting and tracking sensitive information related to the original execution of entities. As shown, embodiments may include or utilize a special-purpose or general-purpose computer system 101, which includes computer hardware such as, for example, one or more processors 102, a system memory 103, one or more data repositories 104, and / or input / output hardware 105.
[0025] Embodiments within the scope of the present invention include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media accessible by the computer system 101. A computer-readable medium that stores computer-executable instructions and / or data structures is a computer storage device. A computer-readable medium that carries computer-executable instructions and / or data structures is a transmission medium. Thus, by way of example and not limitation, embodiments of the present invention can include at least two distinctly different kinds of computer-readable media: computer storage devices and transmission media.
[0026] A computer storage device is a physical hardware device that stores computer-executable instructions and / or data structures. Computer storage devices include various computer hardware such as RAM, ROM, EEPROM, solid-state drives (“SSDs”), flash memory, phase change memory (“PCM”), optical disc storage, magnetic disk storage, or any other hardware device that can be used to store program code in the form of computer-executable instructions or data structures and can be accessed and executed by the computer system 101 to implement the functions disclosed in the present invention. Thus, for example, a computer storage device can include the illustrated system memory 103, the illustrated data repository 104 that can store computer-executable instructions and / or data structures, or other storage means such as on-processor storage described later.
[0027] A transmission medium can include a network and / or a data link that can be used to carry program code in the form of computer-executable instructions or data structures and can be accessed by the computer system 101. A “network” is defined as one or more data links that enable the transmission of electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to the computer system via a network or another communication connection (wired, wireless, or a combination of wired or wireless), the computer system can consider the connection to be a transmission medium. Combinations of the above should also be included within the scope of computer-readable media. For example, the input / output hardware 105 can include hardware (e.g., a network interface module (e.g., “NIC”)) that connects a network and / or a data link that can be used to carry program code in the form of computer-executable instructions or data structures.
[0028] In addition, after reaching various computer system components, program code in the form of computer-executable instructions or data structures can be automatically transferred from a transmission medium to a computer storage device (and vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a NIC (e.g., input / output hardware 105) and then ultimately transferred to system memory 103 and / or a less volatile computer storage device (e.g., data repository 104) at the computer system 101. Thus, it should be understood that a computer storage device can be included in computer system components that also (or even primarily) utilize a transmission medium.
[0029] For example, computer-executable instructions include instructions and data that, when executed at the processor 102, cause the computer system 101 to perform a particular function or group of functions. The computer-executable instructions can be, for example, binary, intermediate format instructions (such as assembly language) or even source code.
[0030] Those skilled in the art will appreciate that the present invention can be practiced in a network computing environment having many types of computer system configurations, including personal computers, desktop computers, laptop computers, messaging processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, and the like. The present invention can also be practiced in a distributed system environment where local and remote computer systems, linked by a network (either by a hardwired data link, a wireless data link, or a combination of hardwired and wireless data links), both execute tasks. Thus, in a distributed system environment, a computer system can include multiple constituent computer systems. In a distributed system environment, program modules can be located in both local and remote memory storage devices.
[0031] Those skilled in the art should also appreciate that the present invention can be practiced in a cloud computing environment. A cloud computing environment can be distributed, although this is not required. When distributed, a cloud computing environment can be distributed internationally within an organization and / or have components that are owned across multiple organizations. In this specification and the appended claims, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services). The definition of "cloud computing" is not limited to any of the many other advantages that can be obtained from such a model when appropriately deployed.
[0032] A cloud computing model can consist of various characteristics such as on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, etc. The cloud computing model can also take the form of various service models such as, for example, Software as a Service (“SaaS”), Platform as a Service (“PaaS”), and Infrastructure as a Service (“IaaS”). The cloud computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, etc.
[0033] Some embodiments, such as cloud computing environments, can include a system that includes one or more hosts, each of which is capable of running one or more virtual machines. During operation, the virtual machines simulate an operating computing system, support an operating system, and perhaps also one or more other applications. In some embodiments, each host includes a hypervisor that simulates the virtual resources of the virtual machines using physical resources that are abstracted from the perspective of the virtual machines. The hypervisor also provides appropriate isolation between the virtual machines. Thus, from the perspective of any given virtual machine, the hypervisor provides the illusion that the virtual machine is interfacing with the physical resources, even though the virtual machine only interacts with the appearance of the physical resources (e.g., virtual resources). Examples of physical resources include processing power, memory, disk space, network bandwidth, media drives, etc.
[0034] Figure 1 A simplified representation of the internal hardware components including processor 102 is shown. As shown, each processor 102 includes a plurality of processing units 102a. Each processing unit can be physical (i.e., a physical processor core) and / or logical (i.e., a logical core provided by a physical core that supports hyper-threading, where more than one application thread is executed at the physical core). Thus, for example, even if processor 102 can include only a single physical processing unit (core) in some embodiments, it can include two or more logical processing units 102a presented by that single physical processing unit.
[0035] Each processing unit 102a executes processor instructions defined by an application (e.g., tracker 104a, operation kernel 104f, application 104g, etc.), and these instructions are selected from a predefined processor instruction set architecture (ISA). The specific ISA of each processor 102 varies based on the processor manufacturer and processor model. Common ISAs include the IA-64 and IA-32 architectures of INTEL, INC., the AMD64 architecture of ADVANCED MICRO DEVICES, INC., and various Advanced RISC Machines (“ARM”) architectures of ARM HOLDINGS, PLC, but there are a large number of other ISAs and these ISAs can be used by the present invention. Generally, an “instruction” is the smallest externally visible (i.e., outside the processor) code unit that is executable by the processor.
[0036] Each processing unit 102a fetches processor instructions from one or more processor caches 102b and executes the processor instructions based on the data in the (multiple) caches 102b, based on the data in the registers 102d, and / or in the absence of input data. Generally, each cache 102b is a small amount (i.e., small relative to the typical amount of system memory 103) of random access memory that stores on-processor copies of portions of the backing store, such as system memory 103 and / or another cache in the cache 102b. For example, when executing application code 103a, one or more caches 102b contain portions of the application runtime data 103b. If the processing unit 102a requests data that is not already stored in a particular cache 102b, a "cache miss" occurs and the data is fetched from the system memory 103 or another cache, potentially "evicting" some other data from that cache 102b. The cache 102b can include a code cache portion and a data cache portion. When executing application code 103a, the code portion of the (multiple) caches 102b can store at least a portion of the processor instructions stored in the application code 103a, and the data portion of the cache 102b can store at least a portion of the data structures of the application runtime data 103b.
[0037] Each processor 102 also includes microcode 102c, which includes control logic (i.e., executable instructions) that controls the operation of the processor 102 and generally serves as an interpreter between the hardware of the processor and the processor ISA exposed by the processor 102 to execute applications. The microcode 102 is typically implemented in a processor-based storage device such as ROM, EEPROM, etc.
[0038] The registers 102d are hardware-based storage locations that are defined based on the ISA of the (multiple) processors 102 and from which values are read or written via processor instructions. For example, the registers 102d are typically used to store values fetched from the cache 102b for use by instructions, store the results of executed instructions, and / or store status or state, such as certain side effects of executed instructions (e.g., sign change of a value, value reaching zero, carry occurring, etc.), processor cycle count, etc. Thus, some of the registers 102d can include "flags" that are used to signal a certain state change caused by the execution of a processor instruction. In some embodiments, the processor 102 can also include control registers for controlling different aspects of the processor operation. Although Figure 1Register 102d is shown as a single box, but it should be understood that each processing unit 102a typically includes one or more corresponding sets of registers 102d specific to that processing unit.
[0039] Data repository 104 may store computer-executable instructions representing an application, such as, for example, tracker 104a, indexer 104b, debugger 104c, security component 104d, operating system kernel 104f, application 104g (e.g., the application that is the subject of the tracking performed by tracker 104a). When these programs are executing (e.g., using processor 102), system memory 103 may store corresponding runtime data, such as runtime data structures, computer-executable instructions, etc. Thus, Figure 1 System memory 103 is shown as including time application code 103a and application runtime data 103b (e.g., each corresponding to application 104g). Data repository 104 may also store data structures, such as tracking data stored in one or more tracking data repositories 104e. As shown by ellipse 104h, data repository 104 may also store other computer-executable instructions and / or data structures.
[0040] Tracker 104a may be used to record a bit-exact trace of the execution of one or more entities, such as one or more threads of application 104g or kernel 104f, and store the tracking data in tracking data repository 104e. In some embodiments, tracker 104a is a stand-alone application, while in other embodiments, tracker 104a is integrated into another software component, such as kernel 104f, hypervisor, cloud fabric, etc. Although tracking data repository 104e is shown as part of data repository 104, tracking data repository 104e may also be implemented at least partially in system memory 103, in cache 102b, or at some other storage device.
[0041] In some embodiments, tracker 104a records a bit-exact trace of the execution of one or more entities. As used herein, a "bit-exact" trace is a trace that includes enough data to support the code that was previously executed at one or more processing units 102a to be replayed such that it executes in substantially the same manner at replay time as it did during the trace. There are a variety of methods that tracker 104a may use to record a bit-exact trace. Two different families of methods that provide a high level of performance and a reasonable trace size will now be briefly outlined, but it should be understood that embodiments herein may operate with traces recorded using other methods. Additionally, optimizations may be applied to any of these families of methods that are not described herein for the sake of brevity.
[0042] The first family of methods is based on the recognition that processor instructions (including virtual machine "virtual processor" instructions) generally fall into one of the following three categories: (1) instructions identified as "non-deterministic" because they do not produce a predictable output since their output is not fully determined by the data in the general-purpose registers 102d or the cache(s) 102b, (2) deterministic instructions whose inputs do not depend on memory values (e.g., they depend only on processor register values or values defined in the code itself), and (3) deterministic instructions whose inputs depend on values read from memory. Thus, in some embodiments, storing sufficient state data to reproduce the execution of an instruction can be achieved by addressing the following problems: (1) how to record non-deterministic instructions that produce outputs not fully determined by their inputs, (2) how to reproduce the values of the input registers for an instruction depending on the registers, and (3) how to reproduce the values of the input memory for an instruction depending on the memory reads.
[0043] In some embodiments, the (multiple) first methods for recording a trace record non-deterministic instructions that produce outputs not fully determined by their inputs by storing the side effects of the execution of such instructions into the trace data repository 104e. As used herein, a "non-deterministic" instruction can include somewhat less common instructions that (i) produce non-deterministic outputs each time they are executed (e.g., RDTSC on an INTEL processor that writes the number of processor cycles since the last processor reset into a register), (ii) can produce deterministic outputs but depend on inputs not traced by the tracer 104a (e.g., debug registers, timers, etc.), and / or (iii) produce processor-specific information (e.g., CPUID on an INTEL processor that writes processor-specific data into a register). Storing the side effects of the execution of such instructions can include, for example, storing the register values and / or memory values changed by the execution of the instruction. In some architectures such as those from INTEL, processor features such as those found in the virtual machine extensions (VMX) can be used to capture instructions for recording their side effects into the trace data repository 104e.
[0044] Solving how to reproduce the values of input registers for deterministic instructions (e.g., whose inputs depend only on processor register values) is straightforward because they are the outputs of the execution of previous instruction(s). Thus, the (multiple) first method for recording a trace can therefore reduce the recording of the execution of an entire series of processor instructions into the trace data repository 104e to reproduce register values at the start of the series; the trace data in the trace data repository 104e does not need to store a record of which specific instructions are executed in the series, or intermediate register values. This is because the actual instructions are available from the application code 103a itself. Thus, the recorded inputs (i.e., the initial set of recorded register values) can be provided to these instructions during replay to execute them in the same manner as they were during the trace.
[0045] Finally, the (multiple) first method for recording a trace can solve how to reproduce the input memory values for deterministic instructions whose inputs depend on memory values by recording the memory values (i.e., reads) consumed by these instructions into the trace data repository 104e, regardless of how the values read by the instructions are written to memory. In other words, some embodiments include recording only memory reads and not memory writes. For example, although values can be written to memory by the current thread, by another thread (including the kernel, e.g., as part of handling an interruption) or by a hardware device (e.g., input / output hardware 105), it is only the values of the instruction reads of the thread that are required to fully replay the instructions of the thread that executed the reads. This is because the values read by the thread (not necessarily all the values written to memory) determine how the thread executes.
[0046] A second family of methods for recording bit-exact traces is founded on the recognition that the processor 102 (including the cache 102b) forms a semi-closed or quasi-closed system. For example, once a portion of the data for a process (i.e., code data and runtime application data) is loaded into the cache(s) 102b, the processor 102 can operate on its own as a semi-closed or quasi-closed system for multiple bursts of time without any input. In particular, once the cache(s) 102b is / are loaded with data, one or more processing units in the processing unit 102a use the runtime data stored in the data portion(s) of the cache(s) 102b and use the registers 102d to execute the instructions from the code portion(s) of the cache(s) 102b. When a processing unit 102a needs some influx of information (e.g., because an instruction it is executing, about to execute, or could execute accesses code or runtime data not in the cache 102b), a "cache miss" will occur and the information is brought into the cache 102b from the system memory 103. For example, if a data cache miss occurs when an instruction performs a memory operation at a memory address within the application runtime data 103b, the data from that memory address is brought into one of the cache lines of the data portion of the cache 102b. Similarly, if a code cache miss occurs when an instruction performs a memory operation at a memory address of the application code 103a stored in the system memory 103, the code from that memory address will be brought into one of the cache lines of the code portion of the cache 102b. Then, the processing unit 102a continues execution using the new information in the cache 102b until the new information is brought into the cache 102b again (e.g., due to another cache miss or a non-cached read).
[0047] Accordingly, in the second family of methods, the tracker 104a can record enough data so that the influx of information can be reproduced into the cache 102b when the traced processing unit executes. Four example implementations within this second family of methods will now be described, although it should be understood that these are not exhaustive.
[0048] A first implementation can record all cache misses and non-cached reads (i.e., reads from hardware components and non-cacheable reads), and the time at which each piece of data is brought into cache 102b during execution (e.g., using a count of the executed instructions or some other counter) to record all data that will be brought into cache 102b in the trace data repository 104e. Thus, the effect is to record a log of all data consumed by the processing unit 102a being traced during code execution. However, due to the alternating execution of multiple threads and / or speculative execution, this implementation may record more data than is necessary to replay the execution of the traced code.
[0049] A second implementation in the second family of methods improves upon the first implementation by only tracking and recording cache lines "consumed" by each processing unit 102a, and / or only tracking and recording a subset of the cache lines used by the processing units 102a participating in the tracing, rather than recording all cache misses. As used herein, a processing unit "consumes" a cache line when it knows the current value of the cache line. This may be because the processing unit is the unit that wrote the current value of the cache line, or because the processing unit performed a read of the cache line. Some embodiments track the consumed cache lines along with extensions to one or more caches 102b (e.g., additional "logging" or "accounting" bits) that enable the processor 102 to identify for each cache line one or more processing units 102a that consumed the cache line. Embodiments can track a subset of the cache lines being used by the processing units 102a participating in the tracing by using way locking in an associative cache. For example, the processor 102 can dedicate a subset of the ways in each address group of the associative cache to the traced processing units and only log cache misses related to those ways.
[0050] A third implementation in the second family of methods can additionally or alternatively be built on top of the cache coherence protocol (CCP) used by the cache(s) 102b. In particular, the third implementation can use the CCP to determine a subset of "consumed" cache lines to record in the trace data repository 104e, and this will still enable the activity of the cache(s) 102b to be reproduced. This method can operate at a single cache level (e.g., L1) and log data flushes and the log of CCP operations at the granularity of the processing unit that caused the given CCP operation to that cache level. This includes logging which processing unit(s) have previously read and / or written access to the cache line.
[0051] A fourth implementation may also utilize CCP data, but operates when two or more cache - level log records of data going to a "higher - level" shared cache (e.g., in the L2 cache) are incoming, and uses a CCP of at least one "lower - level" cache (e.g., one CCP and another L1 cache) to log a subset of the CCP state transitions for memory locations of each cache (i.e., between segments of "load" operations and segments of "store" operations). The effect is that less CCP data is logged compared to the third implementation (i.e., because it logs much less CCP state data than the third implementation, as it is based on load / store transitions rather than per - processing - unit activity). Such logs can be post - processed and enhanced to achieve the level of detail recorded in the third implementation, but can also potentially be built into silicon with lower - cost hardware modifications than the third implementation (e.g., because less CCP data needs to be tracked and recorded by the processor 102).
[0052] Regardless of the recording method used by the tracker 104a, it can record the trace data into one or more trace data repositories 104e. As an example, the trace data repository 104e can include one or more trace files, one or more regions of physical memory, one or more regions of a processor cache (e.g., L2 or L3 cache), or any combination or multiple of these. The trace data repository 104e can include one or more trace data streams. In some embodiments, for example, multiple entities (e.g., processes, threads, etc.) can each be traced to a separate trace file or trace data stream within a given trace file. Alternatively, data packets corresponding to each entity can be marked such that they are identified as corresponding to that entity. If multiple related entities (e.g., multiple threads of the same process) are being traced, then the trace data for each entity can also be traced independently (so that they can be replayed independently), although any event that can be ordered across entities (e.g., accesses to shared memory) can be identified using a global sequence number (e.g., a monotonically increasing number) that is across the independent traces. The trace data repository 104e can be configured for flexible management, modification, and / or creation of trace data streams. For example, modifying an existing trace data stream may involve modifying an existing trace file, replacing segments of trace data within an existing file, and / or creating a new trace file that includes the modification.
[0053] In some implementations, the tracker 104a can be continuously appended to the (multiple) tracking data streams such that the tracking data grows continuously during tracking. However, in other implementations, the tracking data streams can be implemented as one or more circular buffers. In such implementations, as new tracking data is added to the tracking data repository 104e, the oldest tracking data is deleted from the data stream. Thus, when the tracking data streams are implemented as buffers, they contain a rolling trace of the most recent execution at the process being traced. The use of circular buffers can support the tracker 104a to perform "always-on" tracing even in production systems. In some implementations, tracing can actually be enabled and disabled at any time. Thus, whether traced to a circular buffer or appended to a traditional tracking data stream, the tracking data can include intervals between periods during which the tracing was enabled.
[0054] The tracking data repository 104e can include information that helps facilitate efficient trace replay and search on the tracking data. For example, the tracking data can include periodic key frames that support replaying the tracking data stream starting from the moment of the key frame. The key frames can include, for example, the values of all processor registers 102d required to resume the replay. The tracking data can also include memory snapshots (e.g., the values of one or more memory addresses at a given time), reverse lookup data structures (e.g., identifying information in the tracking data based on a memory address as a key), etc.
[0055] Even when using the effective tracing mechanisms described above, there can be practical limitations on the richness of the information that can be stored in the tracking data repository 104e during tracing by the tracker 104a. This can be due to efforts to reduce memory usage, processor usage, and / or input / output bandwidth usage during tracing (i.e., reducing the impact of tracing on the application being traced), and / or reducing the amount of generated tracking data (i.e., reducing disk space usage). Thus, even though the tracking data can include rich information such as key frames, memory snapshots, and / or reverse lookup data structures, the tracker 104a can still limit the frequency of recording this information to the tracking data repository 104e, or can even omit some of these types of information entirely.
[0056] To overcome these limitations, embodiments can include an indexer 104b that takes as input the trace data generated by a tracer 104a and performs a transformation on the trace data to improve the consumption performance of a debugger 104c with respect to the trace data (or data derived therefrom). For example, the indexer 104b can add key frames, memory snapshots, reverse lookup data structures, etc. The indexer 104b can augment existing trace data and / or can generate new trace data containing new information. The indexer 104b can operate based on a static analysis of the trace data and / or can perform runtime analysis (e.g., based on replaying one or more portions of the trace data).
[0057] The debugger 104c can be used to consume (e.g., replay) the trace data generated by the tracer 104a into a trace data repository 104e, including any derivation of the trace data generated by the indexer 104b (either concurrently or on other computer systems), to assist a user in performing debugging operations on the trace data (or data derived therefrom). For example, the debugger 104c can present one or more debugging interfaces (e.g., user interfaces and / or application programming interfaces), replay prior to execution of one or more portions of an application 104g, set breakpoints / watchpoints (including reverse breakpoints / watchpoints), enable query / search of the trace data, and so on.
[0058] A security component 104d identifies sensitive information (i.e., data and / or code) captured by the tracer 104a and takes one or more actions to ensure that such information is restricted and not presented at the debugger 104c. With respect to sensitive data, this can include one or more of the following: preventing sensitive data from being placed in the trace data repository 104e, deleting sensitive data from the trace data repository 104e, masking / encrypting sensitive data in the trace data repository 104e, segregating sensitive data in the trace data repository 104e (e.g., by storing it in a separate trace data stream), modifying the trace data such that execution during replay of the trace data is modified to avoid presenting sensitive data, modifying the trace data such that the execution path taken during replay of the trace data is the same as the path taken during tracing, even if the modified trace data lacks sensitive data, to prevent the debugger 104c from presenting sensitive data even if the sensitive data is present in the unmodified trace data, and so on. With respect to sensitive code, this can include deleting code from the trace data, bypassing code in the trace data 104e, encrypting code in the trace data 104e, and so on. In conjunction Figure 2 Example embodiments of the security component 104d are described in more detail.
[0059] In some implementations, the security component 104d enhances the functionality of one or more of the tracker 104a, the indexer 104b, and the debugger 104c. Thus, for example, the security component 104d can enhance the tracker 104a to have the ability to avoid writing sensitive information to the trace data repository 104e and / or protect sensitive information in the trace data repository 104e; the security component 104d can enhance the indexer 102b to have the ability to purge trace data of sensitive information and / or mask sensitive information in the trace data; and / or the security component 104d can enhance the debugger 104c to have the ability to avoid presenting sensitive information contained in the trace data.
[0060] Although the tracker 104a, the indexer 104b, the debugger 104c, and the security component 104d (for clarity) are depicted as separate entities, it should be understood that one or more of these entities can be combined (e.g., sub-components) into a single entity. For example, a debugging suite can include each of the tracker 104a, the indexer 104b, the debugger 104c, and the security component 104d. In another example, a tracing suite can include the tracker 104a and the indexer 104b, while the debugging suite can include the debugger 104c; alternatively, the tracing suite can include the tracker 104a, while the debugging suite can include the indexer 104b and the debugger 104c. In these latter examples, the security component 104d can be implemented in each of the tracing suite and the debugging suite, or can be implemented as a common library shared by these suites. Of course, there can also be other variations. Notably, the tracker 104a, the indexer 104b, the debugger 104c, and the security component 104d do not need to all be present at the same computer system. For example, the tracing suite can execute at one or more first computer systems (e.g., production environment, test environment, etc.), while the debugging suite can execute at one or more second computer systems (e.g., a developer's computer, a distributed computing system that facilitates distributed replay of trace data, etc.). Also, as shown, the tracker 104a, the indexer 104b, and / or the debugger 104c can access the trace data repository 104e directly (i.e., as shown by the dashed arrow) and / or through the security component 104c (i.e., as shown by the solid arrow).
[0061] As described above, Figure 2 is shown such as Figure 1Example embodiments of security components such as security component 104d of. As shown, security component 200 may include a plurality of sub-components, such as, for example, identification component 201 (including annotation sub-component 201a, derivation sub-component 201b, replication sub-component 201c, user input sub-component 201d, database 201e, etc.), data modification component 202, code modification component 203, etc. Although these components are presented as aids in describing the functionality of security component 200, it should be understood that the specific number and identity of these components may vary according to the implementation.
[0062] Identification component 201 identifies sensitive original information. As previously mentioned, this may include potentially sensitive information and definitely sensitive information. The original information may correspond to any form of code or data accessed by processor 102 and is typically stored at one or more memory addresses. The original information may correspond to at least a portion of one or more of a pointer, data structure, variable, class, field, function, source file, component, module, executable instruction, etc.
[0063] The specific time at which identification component 201 identifies sensitive information may vary depending on the implementation. For example, identification may occur during the initial recording into data repository 104e by tracker 104a (e.g., runtime analysis such as using page table entries, enclave region metadata, etc.), during the post-processing of the trace data by indexer 104b (e.g., static analysis that may use debug symbols, etc.), and / or during the replay of the trace data by indexer 104b and / or debugger 104c (e.g., static and / or runtime analysis). When it identifies sensitive information, identification component 201 may record it in one or more databases 201e for use by replication component 102c, as described later.
[0064] Identification component 201 may identify sensitive information in various ways represented by annotation sub-component 201a, derivation sub-component 201b, replication sub-component 201c, and user input sub-component 201d (although other methods may also be used, as indicated by the ellipsis). Annotation sub-component 201a, derivation sub-component 201b, and replication sub-component 201c will now be described in detail. However, initially it should be noted that user input sub-component 201d may be used in conjunction with any of these components, such as to manually identify code that interacts with sensitive data (i.e., as input to annotation sub-component 201a), manually identify data derived from data (i.e., as input to derivation sub-component 201b), manually identify a copy of data (i.e., as input to replication sub-component 201c), and / or provide data for database 201e.
[0065] Typically, the annotation sub-component 201a can identify sensitive information based on annotations about an entity being executed at the processor 102 and being tracked. For example, the code of an entity (whether source code, object code, assembly code, machine code, etc.) can be annotated to identify one or more portions of the code that are sensitive in themselves, which take sensitive data as input, process sensitive data, generate sensitive data, store sensitive data, and so on. For example, an entity can be annotated to identify functions, parameters, modules, variables, data structures, source code files, input fields, etc. that are sensitive in themselves and / or may be involved in the creation, consumption, or processing of sensitive data. As an example, a memory region of a secure enclave can be considered sensitive (for code, data reads, and / or data writes). As an example, page table entries applicable during original execution can indicate permissions applied to a portion (or all) of the trace stream. As an example, PCID, ASID, processor security level, processor exception level, etc. can be used to determine security boundaries applicable to a portion (or all) of the trace stream. Thus, the annotation sub-component 201a can use these annotations to identify when sensitive code portions are executed (e.g., as part of tracing the code, replaying in the code, and / or analyzing a portion of the code based on its previous execution being traced to the data repository 104e), and / or to access, process, and / or generate sensitive data when the code portion is executed.
[0066] The annotations on which the annotation sub-component 201a depends can be added to the code of the entity itself, can be added as a separate metadata file, and / or can even be stored in the database 201e. These annotations can be created in various ways, such as using manual, automatic, and / or semi-automatic techniques, such as manual input (i.e., using the user input component 201d), machine learning, derived analysis performed by the annotation sub-component 201a, etc. Annotations about an entity can be created before and / or during the analysis performed by the annotation sub-component 201a (e.g., based on user input, machine learning, derived analysis, etc.).
[0067] On the other hand, the derivation sub-component 201b utilizes the rich nature of time-travel tracing (i.e., the fact that they capture bit-exact traces of the way previous code was executed) to trace code execution, including data flow during code execution, in one or both of the forward and backward directions on the execution timeline. As part of this tracing, the derivation sub-component 201b can identify at least one or more of the following: (i) data derived from data that has been identified as sensitive; (ii) data from which data that has been identified as sensitive itself is derived; (iii) code that may be annotated as related to sensitive data because it acts on data previously acted upon by code known to be related to data identified as sensitive; (iv) code that can be annotated as related to sensitive data because it acts on data later acted upon by code known to be related to data identified as sensitive; or (v) code that can itself be annotated as sensitive because it has execution continuity with code that has been identified as sensitive. To further understand these concepts, Figure 3A and 3B illustrates an example of derivation analysis for identifying derived data and / or code.
[0068] Figure 3A Illustrates an example embodiment 300a that relies on bit-exact tracing to identify derived data and / or code in the forward direction on the execution timeline. Example 300a includes a portion of a timeline 301 representing execution at a processing unit 102a, and a corresponding portion of a traced data stream 302 that stores the bit-exact trace of that execution. Example 300a can represent tracing the original execution of an entity into the traced data stream 302, or replaying the original execution of the entity from the traced data stream 302. Figure 3A Illustrates two points 303 and 305, both representing different execution moments during the timeline 301, and corresponding portions of data into the traced data stream 302. In particular, point 303 represents a specific moment at which first data being accessed by a first code has been identified as sensitive (and by extension, the first code acts on sensitive data) and / or the first code itself is known to be identified as sensitive data. On the other hand, point 305 represents a later execution moment at which second data is being accessed by a second code and / or the second code itself is being accessed or executed.
[0069] At point 305, it is not yet known whether the second data is sensitive; whether the second code acts on sensitive data; and / or whether the second code itself is sensitive. However, Figure 3AAn arrow 304 from point 303 to point 305 is also shown. Arrow 304 indicates that there is a traced code and / or data continuity between points 303 and 305, which can be used to determine that the second data accessed at point 305 can be identified as sensitive because it is derived from the first data accessed at point 303 and / or the second code accessed / executed at point 305 itself can also be sensitive due to its relationship with the first code executed at point 303. Thus, the continuity represented by arrow 304 enables the derivation sub-component 201b to analyze / replay the traced data stream 302 in the forward direction (i.e., from point 303 to point 305), and determine that the second data accessed at point 305 is derived from the first data accessed at point 303 and / or the second code is executed as a result of the first code.
[0070] Based on this continuity, the derivation sub-component 201b can then identify the second data as sensitive data; the second code acts on sensitive data; and / or the second code itself is sensitive. If the second code itself is identified as acting on sensitive data and / or being sensitive data, the annotation sub-component 201a can annotate the second code accordingly (if needed). Although only one instance of the derived data and / or code is shown in Figure 3A it should be understood that this analysis can be applied recursively to identify any number of additional derived items.
[0071] On the other hand, Figure 3B an example 300b is shown that relies on bit-exact tracing to identify derived data and / or code in the opposite direction on the execution timeline. Similar to Figure 3A Figure 3B it shows a portion of the timeline 306 representing the execution at the processing unit 102a, and the corresponding portion of the traced data stream 307 that stores the bit-exact trace of the execution. Figure 3B It also includes two points 308 and 310, both representing different execution moments during the timeline 306, and the corresponding portions of the data in the traced data stream 307. However, different from Figure 3A Figure 3BAn arrow 309 is shown, in the opposite direction (i.e., from point 310 to point 308), and the arrow 309 represents the traced code / data continuity between these points. This indicates that point 310 is the point where the accessed data is identified as sensitive and / or the executed code is identified as sensitive; and this point 308 is the point where the relevant data and / or code is accessed (i.e., a previous time during the execution timeline 306). Here, the continuity represented by arrow 309 enables the derivation sub-component 201b to determine, in the opposite direction (i.e., from point 310 to point 308), that the data accessed at point 310 is derived from the data accessed at point 308 and / or the code executed at point 310 is related to the code executed at point 308. Thus, the derivation sub-component 201b can identify the data accessed at point 308 as sensitive data; the code executed at point 308 interacts with sensitive data; and / or the code executed at point 308 itself is sensitive code. Again, the annotation sub-component 201a can annotate the code as needed and / or this analysis can be applied recursively (forward from point 310 and / or backward from point 308) to identify any number of additional derived items. In an implementation, this allows sensitive data to be identified (and thus protected) backward in execution time.
[0072] The copy sub-component 201c identifies copies of sensitive information based on the database 201e, which identifies or includes information that has previously been identified as sensitive. The copy sub-component 201c can make these identifications at runtime (e.g., as part of the tracing using the tracker 104a, and / or as part of the replay of the indexer 104b and / or the debugger 104c) or statically (e.g., as part of a static analysis of the trace data such as performed by the indexer 104b). Generally, the copy sub-component 201c can determine whether items accessed, generated, and / or executed at runtime (either during tracing or during replay) and / or items that have been logged into the trace data repository 104e can be identified as sensitive items by directly comparing them with the entries in the database 201e and / or by comparing some of their derived items with the entries in the database 201e.
[0073] As described above, the database 201e can be generated by the identification component 201 and / or can be provided to the identification component 201 in some other way (e.g., by a human user). For example, when the identification component 201 identifies an item as a sensitive item (e.g., based on the operation of the annotation component 201a, the operation of the derivation component 201b, or user input at the user input component 201d), the identification component 201 can store the item itself or its derived item in a form of searchable data repository (such as a lookup table, a hash table, an SQL data table, etc.). In this way, in operation, the identification component 201 can develop an evolving knowledge base of the identified sensitive information, which the copy sub-component 201c can rely on. In some implementations, data type information (e.g., from debug symbols, function prototype information, etc.) can enhance the identification of data, such as by indicating how sensitive data is embodied within a structure, a string, or other data type. As described above, the database 201e can store the information itself or its derived information. Storing the derived information of the information (e.g., the hash of the information, the encrypted version of the information, etc.) can provide multiple benefits, such as maintaining the security of the information, reducing the storage requirements in the database 201e, and / or improving the storage and / or search performance of the database 201e.
[0074] When the task of the copy sub-component 201c is to determine whether the original information is sensitive, it can compare the original information with the entries in the database(s) 201e to determine whether it has previously been identified as sensitive information in the database(s) 201e. If so, the copy sub-component 201c can also identify the information instance as sensitive information. If the database 201e stores the hash or encrypted version of the sensitive information, the copy sub-component 201c can hash or encrypt the original information using the same algorithm / key before the comparison.
[0075] By maintaining and using the searchable (multiple) database 201e in this manner, the identification component 201 (utilizing the replicating sub-component 201c) can help identify sensitive information that may not typically be identified by the annotating sub-component 201a and / or the deriving sub-component 201b in certain situations. The reason is that even if the annotating sub-component 201a and / or the deriving sub-component 201b can identify an item as a sensitive item at a first moment during code execution (e.g., because it is used in conjunction with code known to interact with sensitive data), the same information may not be identified as sensitive by these components at a second moment during code execution (regardless of whether the second moment is before or after the first moment). Thus, in at least some implementations, an earlier access to data in the bit-precise tracking can be determined to be sensitive based on a subsequent (later in execution time) access to the data, thereby identifying the data as sensitive data. In at least some cases, this allows for the protection of information even at first use, even when that first use itself does not result in the data being identified as sensitive data, and prior to any access (execution time) that results in the data being identified as sensitive data.
[0076] For example, even if a data item or code is in fact related to sensitive information already identified in the database 201e through code continuity, that code continuity may not be captured in the tracking data repository 104e. This can occur, for example, in cases where the tracking data repository 104e is missing tracking data during one or more execution periods. The tracking data repository 104e may be missing such data if tracking is enabled and disabled while being recorded and / or if the tracking data stream is implemented as a circular buffer.
[0077] In cases where tracking is enabled / disabled while being recorded, sensitive data can be read by a first code while tracking is enabled, tracking can subsequently be disabled, and when tracking is subsequently enabled, the same data can be read again by a second (but related) code. In such a case, based on the annotating sub-component 201a knowing that the first code interacts with sensitive data, the sensitive data can be identified in relation to the first read; however, the annotating sub-component 201a may not be able to identify the same data as sensitive during a subsequent tracking period because the annotating sub-component 201a lacks knowledge of the second code interacting with the sensitive data. Additionally, due to gaps in the available tracking data, the deriving sub-component 201b may not be able to track the continuity between the first read and the second read. However, the replicating sub-component 201c can recognize that the second code is reading data known to be sensitive and thus identify the second read as sensitive. This principle also applies in the other direction, i.e., based on the second read being identified as sensitive by the annotating sub-component 201a and / or the deriving sub-component 201b, the replicating sub-component 201c can identify the first read as sensitive.
[0078] In the case of using (multiple) data streams implemented as (multiple) circular buffers, the trace data repository 104e may store one or more memory snapshots that capture values obtained while the trace is active but written to memory when the trace data is unavailable (e.g., because the trace is disabled, or the trace data has been purged from the circular buffer). In these cases, the memory snapshots may contain sensitive data, but the annotation sub-component 201 and / or the derivation sub-component 201b cannot identify this data as sensitive because they do not know what code placed these values in memory (and whether there is continuity between that code and code known to interact with sensitive data). Here, the replication sub-component 201c can compare these values with the database 201e to determine whether any values should be identified as sensitive data. Similarly, this principle also applies in the other direction, i.e., the replication sub-component 201c can identify sensitive data in memory snapshots obtained after the period of the traced execution.
[0079] Even in cases where a copy is completely unrelated to known sensitive information by code continuity, the replication sub-component 201c can identify one or more copies of sensitive information. For example, the same data may be provided as input to (or even generated by) completely independent and unrelated code execution paths. If the data is identified as sensitive in one code execution path, the replication sub-component 201c can also identify it as sensitive when used in another code execution path. This principle also applies to the memory snapshot example above, i.e., even if the data in the snapshot may be completely unrelated to data identified as sensitive during code execution, the replication sub-component 201c can still identify it as sensitive data.
[0080] It should be understood that the replication sub-component 201c can process the entire trace data in the trace data repository 104e to locate all information in the trace data identified as sensitive in the database 201e. Thus, once at least one copy of an item is identified as sensitive, the replication sub-component 201c can identify all other instances of that item in the trace data repository 104e, regardless of how those instances entered the trace data repository 104e. Once identified, this enables the security component 200 to delete / mask all instances of sensitive information from the trace data repository 104e. Similarly, if there are reverse index structures in the trace data repository 104e, once a storage location is identified as containing sensitive data, these reverse index structures can be used to quickly identify other accesses to the sensitive data and / or identify when (at execution time) the sensitive data is overwritten by non-sensitive data.
[0081] Non-sensitive data can include data that is explicitly defined as non-sensitive or data that is found to be non-sensitive in a collective manner (e.g., because it is found to be derived from or a copy of data that has been explicitly defined as non-sensitive). The identity of specific information can change from sensitive to non-sensitive and vice versa. For example, in an implementation, sensitive data is considered sensitive until type-specific conditions are met. For example, a type-specific condition for a string ending with zero can be that all bytes of the original string length are covered by non-sensitive data or all bytes of the original string length are covered by zeros. Similarly, a structure (or a higher-level construct such as a class) can have type-specific conditions that indicate that the structure should continue to be considered sensitive until a destructor or other function is called or a field / member has a specific value (in addition to or as an alternative to a default requirement).
[0082] In some embodiments, the identification component 201 includes functionality for processing code that moves or copies data but does not actually consume or process the data it moves / copies. Examples of such code can be known memory copies (e.g., memcpy, memcpy_s, memmove, memmove_s, etc.) and string copies (e.g., strcpy, strncpy, etc.) used in the C programming language. Such functions can move or copy segments of memory that include sensitive information, but these functions do not actually do anything to the data other than move / copy it (i.e., they are neutral with respect to sensitive data).
[0083] In some implementations, the derivation sub-component 201c and / or the copy sub-component 201c can cause such functions to be marked as related to sensitive information, which may (undesirably) cause all code touched by such functions to be identified as sensitive code. To avoid this, embodiments can maintain a list of known functions that are neutral to sensitive data (e.g., in the database 201e). Then, this list can be used to prevent these functions from being identified as related to sensitive information. Although this can cover known functions that are neutral to sensitive data (such as memcpy and stringcpy), it may not cover custom-coded functions that are neutral to sensitive data. Thus, additionally or alternatively, embodiments can detect code that reads and / or writes data but does not make any decisions based on that data (i.e., it just moves / copies around), and avoid identifying such code as being involved with sensitive information. Example implementations may have annotations in the binary code indicating whether a function is (or is not) neutral to sensitive data, annotations in debug symbols, etc. Such detection can allow for special cases where the special case permits the code to perform limited decisions on the data it moves / copies (e.g., such as performing a copy up to but not including the null terminator) while still being considered neutral to sensitive data.
[0084] Based on the identification component 201, information is identified as sensitive information, and the security component 200 uses one or both of the data modification component 202 or the code modification component 203 to delete it from the trace data repository 104e or mask it within the trace data repository 104e.
[0085] The data modification component 202 replaces data items that have been identified as sensitive with alternative data and / or stores these data items in a masked or protected manner in the trace data repository 104e. For example, Figure 4A and 4B illustrates an example embodiment of data item replacement and / or masking in time travel tracing.
[0086] Initially, Figure 4A illustrates example 400a of sensitive data item replacement / masking for a single trace data stream. In particular, Figure 4A illustrates a portion of a timeline 401 representing the execution of an entity at the processing unit 102a, and a corresponding portion of a trace data stream 402 that stores a bit - exact trace of the execution. Figure 4A It also illustrates the identification of a sensitive data item at a point 403 in the execution, and the replacement or masking of this sensitive data at the corresponding point 404 in the trace data stream 402. The replacement data item can include identifying or generating alternative data to store in the trace data stream 402 instead of the original data identified at point 403. This can include the data modification component 202 generating random data, identifying predefined data, generating derived data (e.g., hash) of the original data, etc. In some embodiments, identifying or generating alternative data can include preserving one or more characteristics of the original data, such as reserving the type of the data (e.g., string, integer, floating - point number, etc.), reserving the size of the data (e.g., integer size, string length, etc.), reserving a portion of the data (e.g., only replacing a subset of a string), etc. Masking the data item can include the data modification component 202 encrypting the data item before storing it in the trace data stream 402, encrypting the entire trace data stream 402 or the trace file, etc.
[0087] On the other hand, Figure 4B illustrates example 400b of sensitive data item replacement / masking for multiple trace data streams. In particular, Figure 4B illustrates a portion of a timeline 405 representing the execution of an entity at the processing unit 102a, and corresponding portions of trace data streams 406 and 407 for storing a bit - exact trace of the execution. Similar to Figure 4A , Figure 4B shows the identification of a sensitive data item at a point 408 in the execution. However, instead of replacing or masking the data item in a single trace data stream,Figure 4B illustrates replacing a data item in a first data stream (i.e., tracking point 409 in data stream 406), while storing it in a second data stream in its original or masked form (i.e., tracking point 410 in data stream 407). Replacing the data item in the tracking data stream 406 can include any mechanism for generating or identifying alternative data as described above in connection with Figure 4A ; storing the data item in a masked form can include any masking mechanism as described above in connection with Figure 4A . Then, when sensitive data should be protected, debugger 104c can use the data item from tracking data stream 406, and when sensitive data does not need to be protected, debugger 104c can use the data item from tracking data stream 407 (e.g., depending on the user using debugger 104c, the computer on which debugger 104c operates, whether a decryption key has been provided, etc.). Notably, tracking data stream 407 does not have to be a complete tracking of the execution timeline 405. For example, tracking data stream 406 can be used to store the complete tracking, while tracking data stream 407 can be used to store a subset of the tracking activities, such as the tracking activities related to sensitive information.
[0088] The data modification component 202 can operate at any time during tracking generation or consumption. For example, when timeline 401 / 405 represents the original execution of an entity, and when the tracking data streams 402 / 406 / 407 are original tracking data streams (e.g., as recorded by tracker 104a), the data modification component 202 can operate. In another example, when timeline 401 / 405 represents a replayed execution of an entity (e.g., by indexer 104b and / or debugger 104c), and when the tracking data streams 402 / 406 / 407 are the resulting / indexed tracking data streams, the data modification component 202 can operate. The data modification component 202 can also operate based on, for example, static analysis of the tracking data by indexer 104b.
[0089] The code modification component 203 stores data into the tracking such that the execution path taken by the entity during its original execution will also be taken during replay, even if the data modification component 202 performs data replacement activities; and / or stores data into the tracking such that alternative executable instructions rather than the original executable instructions are executed during replay of the entity. These concepts are described in connection with Figures 5A - 5C .
[0090] Initially, Figure 5A illustrates example 500a, which ensures that the execution path taken by the entity during its original execution will also be taken during replay even with data replacement. In particular, Figure 5AIllustrated is a portion of a timeline 501 representing the execution of an entity at processing unit 102a, and a corresponding portion of a trace data stream 503 that stores a bit-exact trace of that execution. Figure 5A Also shown is the identification of a sensitive data item at point 505 during execution, which sensitive data item may be replaced by data modification component 202 in trace data stream 503. However, Figure 5A shown is that depending on the value of the sensitive data item, point 505 has caused an alternative execution path 502 to occur. For example, the sensitive data item may have been a parameter of a conditional statement in the code. Thus, replacing the sensitive data item by data modification component 202 may cause the alternative execution path 502 to occur during replay, which would result in an incorrect trace replay. To prevent such replay behavior, code modification component 203 may store trace data at point 506 in trace data stream 503 that ensures that the original execution path can be taken during replay even if data replacement occurs. This is indicated by replay timeline 504, which shows the original execution path being taken.
[0091] In some embodiments, code modification component 203 records trace data that includes one or more alternative executable instructions that will cause the original path to be taken even if data replacement occurs. For example, the original instruction may be replaced with an alternative executable instruction that changes the condition. Additionally or alternatively, code modification component 203 may record trace data that includes code comments that cause the original execution path to be taken even if the resulting execution is a conditional instruction during replay. For example, as described above, some trace embodiments record the side effects of non-deterministic instructions such that these instructions can be replayed later. Embodiments may apply the same principle to deterministic conditional instructions, i.e., the desired result of a conditional instruction may be recorded as a "side effect" and this side effect can be used to produce the desired result during replay even if the actual result of the execution is a conditional instruction.
[0092] Sometimes, security component 200 may avoid using code modification component 203 while still ensuring that the correct execution path is taken despite data modification. For example, data modification component 202 may ensure that the alternative data it uses to replace the original data will produce the same result for the condition. For example, if the result of the condition is based on data size, this may be accomplished by ensuring that the alternative data has the same size as the original data (e.g., string length).
[0093] Figure 5B Illustrated is an example 500b of storing data into a single trace data stream that causes alternative executable instructions to be executed during replay of an entity. In particular, Figure 5BIllustrated is a portion of a timeline 508 representing the execution of an entity at processing unit 102a, and a corresponding portion of a trace data stream 509 storing a bit-exact trace of the execution. Figure 5B Also illustrated at segment 511 is the execution of identified sensitive code. As a result of segment 511 being identified as sensitive code, code modification component 203 may record data at point 512 into trace data stream 509, which effectively bypasses segment 511 during replay, but allows the replay to continue as normal (with less sensitive code executed), as shown at point 513 on replay timeline 510. For example, the data at point 512 may include one or more alternative instructions that replace a call to the sensitive code with one or more instructions that establish the state that would have been had as a result of the execution of segment 511 (e.g., register and memory values), and jump to the instruction that immediately follows segment 511. As part of this operation, code modification component 203 may replace any sensitive data in that state with data modification component 202 as needed. Additionally or alternatively, the data at point 512 may include one or more key frames, such as key frames that cause the replay to skip segment 511. Additionally or alternatively, the data at point 512 may include one or more "side effects", such as side effects that cause existing instructions to bypass the call to segment 511. Regardless of how it is implemented, the data stored at point 512 causes segment 511 of the sensitive code to effectively be converted to a "black box" during replay. Thus, the technical effect is that timeline 508 can be replayed while skipping or bypassing segment 511. In some embodiments, code modification component 203 may capture a memory snapshot and / or key frames at the start and / or end of this segment 511 in order to capture the memory and / or register state, and record these snapshots and / or key frames in trace data stream 509.
[0094] Figure 5C Illustrated is an example 500c of storing data into at least one trace data stream that causes alternative executable instructions to be executed during replay of an entity. In particular, Figure 5C Illustrated is a portion of a timeline 514 representing the execution of an entity at processing unit 102a, and corresponding portions of trace data streams 515, 516, and / or 517 that may be used to store a bit-exact trace of the execution. Figure 5C Also illustrated is segment 521 being identified as the execution of sensitive code (or access to sensitive data). In practice, as a result of identifying segment 521 as sensitive, code modification component 203 may record data at point 522 that effectively bypasses segment 521 during replay into trace data stream 515, as described in connection with Figure 5B above. Thus, similar to Figure 5BFor the traced data stream 509, the traced data stream 515 can be used to replay the execution while bypassing the sensitive instructions in block 521, as shown by point 525 on replay timeline 518.
[0095] Additionally or alternatively, in an implementation, the code modification component 203 can record instructions 523 that will produce some (or all) side effects caused by the execution of block 521 into the traced data stream 516. As an example, the instruction 523 can write a final value to the memory modified by the execution of block 521, and / or the instruction 523 ensures that the register state matches that at the end of the execution of block 521. As a specific example, if block 521 corresponds to an instruction that encrypts original data (sensitive data) into an encrypted form (non-sensitive data) using a private key (sensitive data), the original sensitive data can be modified (as described throughout the specification), and the code in block 521 can be replaced with an instruction 523 that writes the final encrypted data. In this specific example, since the replacement instruction 523 re-creates the side effects of the deleted block 519, this replacement can eliminate any need for snapshots or key frames in the traced data stream 516. The traced data stream 516 is then used to replay the effects of the sensitive instructions in block 521 without actually executing the instructions in block 521, as shown by block 526 on replay timeline 519.
[0096] Additionally or alternatively, in an implementation, the code modification component 203 can record the execution of block 521 into a traced data stream 517 that can be encrypted. This is shown by block 524. Thus, with the necessary permissions, the traced data stream 517 can be used to actually replay the execution of the sensitive code in block 521 (as shown by block 527 on replay timeline 527).
[0097] Any combination of the traced data streams 515, 516, or 517 can be recorded and / or used for debugging. For example, when sensitive code should be protected during replay, the debugger 104c can replay instructions at point 522 in the traced data stream 515, and / or can replay instructions at point 523 in the traced data stream 516. When the sensitive data is not protected during replay (e.g., depending on the user using the debugger 104c, the computer on which the debugger 104c operates, whether the decryption key has been provided, etc.), the debugger 104c can replay from block 527 in the traced data stream 520. It is noted that each traced data stream may not need to include a complete trace of the execution timeline 514. For example, the traced data stream 515 can be used to store the complete trace, while the traced data streams 516 and / or 517 can be used to store subsets of the traced activities, such as those related to the execution of sensitive code.
[0098] Like the data modification component 202, the code modification component 203 can operate at any time during trace generation or consumption, whether during the tracing of the tracer 104a, the indexing of the indexer 104b, and / or the debugging of the debugger 104c. Additionally, the code modification component 203 can operate based on runtime analysis and / or static analysis.
[0099] It should be noted that the embodiments herein can cover Figures 4A - 5C any combination and / or repeated application of the examples shown.
[0100] Figure 6 The flowchart of an example method 600 for protecting sensitive information related to the original execution of a traced entity is illustrated. Method 600 will be described with respect to Figure 1 the components and data of the computer architecture 100, Figure 2 the security component 200, and the examples of FIGS. 3 - 5C.
[0101] As shown, method 600 includes an action 601: identifying that the original information accessed during the original execution of the entity includes sensitive information. In some embodiments, action 601 includes identifying that the original information includes sensitive information, where the original information is accessed based on the original execution of one or more original executable instructions of the entity. For example, during the original execution of the entity at one or more processing units 102a, or from the trace data repository 104e, the identification component 201 can use one or more of the annotation sub - component 201a, the derivation sub - component 201b, the copy sub - component 201c, or the user input sub - component 201d to identify sensitive information items. As explained throughout the text, sensitive information items can include sensitive data and / or sensitive code.
[0102] As shown, method 600 can also include an action 602: storing alternative information while ensuring that the entity takes the same execution path during replay. In some embodiments, action 602 includes storing first trace data including alternative information rather than the original information into a first trace data stream based on the original information including sensitive information, while ensuring that the execution path taken by the entity based on the original information will also be taken during the replay of the original execution of the entity using the first trace data stream. For example, once the identification component 201 identifies sensitive data, the data modification component 202 can replace that data in the trace data repository 104e with alternative data (such as in Figure 4A and 4B the trace data streams 407 and 406), and may also store the sensitive data in the trace data repository 104e in a protected form (such as in the trace data stream 407).
[0103] Ensuring that the execution path adopted by an entity based on original information will also be adopted during the replay of the original execution of the traced entity can be achieved by one or both of a data modification component 202 or a code modification component 203. For example, the data modification component 202 can select alternative data that will result in the same conditional evaluation result as the original data. For example, if the condition is based on string length, this can be achieved by replacing a string with a string of equal length. On the other hand, the code modification component 203 can replace one or more original instructions with alternative instructions that bypass or alter the result of the condition (e.g., the instruction at point 512 in the traced data flow 509, or the instruction at point 522 in the traced data flow 515), the code modification component 203 can annotate one or more instructions to override the result during replay, and / or the code modification component 203 can insert one or more key frames that simulate the result during replay.
[0104] As shown, method 600 can also include action 603: causing the alternative instructions to be executed during replay. In some embodiments, action 602 includes storing second trace data into a second trace data flow based on the original information including sensitive information, the second trace data flow causing one or more alternative executable instructions rather than one or more original executable instructions of the entity to be executed during the replay of the original execution of the entity using the second trace data flow. For example, as explained with respect to the trace data flows 509 and 515 in Figure 5B and 5C the code modification component 203 can store trace data (e.g., at points 512 and 522) that causes the code executed by the original entity to be bypassed during replay. This can include, for example, replacing instruction segments with one or more instructions that bypass the block, storing one or more instructions that replicate the side effects of executing the segment, storing at least one memory snapshot associated with the segment, and / or storing at least one key frame associated with the segment.
[0105] Depending on the specific sensitive information identified in action 601, method 600 can include only one of actions 602 and 603, or can include both actions 602 and 603. As shown, if actions 602 and 603 are executed simultaneously, although they can also be executed serially, they can also be executed in parallel. Additionally, as shown by arrow 604, actions 602 and 603 can be executed in cooperation with each other. Additionally, any combination of actions 602 and 603 can be repeatedly applied, and each repetition can be in any order, in parallel, or in cooperation with each other. Similarly, although actions 602 and 603 refer to a first trace data flow and a second trace data flow, it should be understood that these can be the same trace data flow.
[0106] Note that method 600 can be executed during the activity of any one of tracker 104a, indexer 104b, and / or debugger 104c. In this way, method 600 can be executed during one or both of the following: (i) the original execution of the entity, or (ii) the post-processing of the trace (performed by tracker 104b or debugger 104c) after the original execution of the entity. Additionally, method 600 can be executed whenever a potentially sensitive original information item is encountered at any stage of these phases. In this way, method 600 can be repeated multiple times during trace recording, trace indexing, and / or trace debugging.
[0107] As described above, action 601 can include identifying component 201 using derivation sub-component 201b and / or copy sub-component 201c. If derivation sub-component 201b is used, action 601 can include identifying the resulting data generated by the execution of one or more original executable instructions as also including sensitive information. If copy sub-component 201c is used, action 601 can include identifying a copy of the original information in the trace as including sensitive information. In this case, the copy of the original information can exist at an execution time after the first occurrence of the original information in the trace (e.g., as described in conjunction with Figure 3B ), or at an execution time before the first occurrence of the original information in the trace (e.g., in conjunction with Figure 3B ). The copy of the original information and the original information can be related by code continuity, or can be independent in the trace (e.g., separate user input). The copy of the original information can be used to identify the original information as sensitive, or the original information can be used to identify the copy of the original information as sensitive.
[0108] Accordingly, the embodiments herein identify sensitive information related to time travel tracing (during trace recording and / or at a later time), and delete and / or mask such sensitive information in the trace. As described above, the embodiments can include storing alternative data in the trace (instead of the original data identified as sensitive), replacing the original instructions in the trace with alternative instructions that avoid executing sensitive code or account for correct execution due to data replacement, overriding the execution behavior of one or more instructions, and so on. In this way, the embodiments enable time travel traces to be generated and consumed even in a production environment while keeping sensitive information from being leaked.
[0109] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the above-described features or acts, or the order of the above acts. Rather, the described features and acts are disclosed as example forms of implementing the claims.
[0110] Without departing from the spirit or essential characteristics of the present invention, the present invention may be implemented in other specific forms. The described embodiments should be considered illustrative rather than restrictive in all respects. Therefore, the scope of the present invention is indicated by the appended claims rather than the foregoing description. All changes that fall within the equivalent meaning and scope of the claims should be included within its scope.
Claims
1. A method implemented at a computer system, the computer system including one or more processors, the method for protecting sensitive information related to a trace of an original execution of an entity, the method comprises: accessing trace data, the trace data representing the execution of a plurality of executable instructions of the entity; based on the trace data, identifying that original information includes sensitive information, the original information being accessed based on the execution of one or more first executable instructions of the entity; identifying derived information of the original information, the identification being based on the execution of one or more second executable instructions of the entity that operate on the original information to create the derived information; based on the derived information derived from the original information, determining that the derived information also includes sensitive information; and based on the original information including the sensitive information, performing one or more of the following: storing first trace data including alternative information instead of the derived information, while ensuring that the execution path adopted by the entity based on the derived information will also be adopted during replay of the entity; or storing second trace data, the second trace data causing one or more alternative executable instructions instead of the one or more second executable instructions of the entity to be executed during replay of the entity.
2. The method according to claim 1, wherein the method is performed during post-processing of the trace after execution of the entity.
3. The method according to claim 1, further comprises: identifying that a copy of the original information in the trace data includes sensitive information, the copy of the original information existing at an execution time before a first occurrence of the original information in the trace data, and wherein identifying that the original information includes sensitive information is based on identifying that the copy of the original information in the trace data includes sensitive information.
4. The method according to claim 1, further comprises: identifying that a copy of the original information in the trace data includes sensitive information, the copy of the original information existing at an execution time after a first occurrence of the original information in the trace data, and wherein identifying that the original information includes sensitive information is based on identifying that the copy of the original information in the trace data includes sensitive information.
5. The method according to claim 1, further comprises: identifying that a copy of the original information in the trace data includes sensitive information, and wherein the copy of the original information and the original information are not related by code continuity.
6. The method according to claim 1, wherein identifying that the original information includes sensitive information is based on determining that the original information is of a type selected from the list consisting of: a specific data structure, a specific variable, a specific class, a specific field, a specific function, a specific source file, a specific component, a specific module, or an executable instruction.
7. The method according to claim 1, wherein the original information is identified as sensitive until a type-specific condition of the type associated with the original information has been satisfied.
8. The method according to claim 1, wherein the method stores the first trace data, and wherein ensuring that the execution path adopted by the entity based on the derived information will also be adopted during replay of the original execution of the entity using the first trace data includes one or more of: recording side effects of one or more instructions, recording one or more alternative instructions, or ensuring that the alternative information will produce the same conditional evaluation result as the derived information.
9. The method according to claim 1, wherein the method stores the second trace data, and wherein storing the second trace data includes one or more of: replacing the segment with one or more instructions segmented by a bypass instruction, replacing the instruction segment with one or more instructions that replicate the side effects of the executed instruction segment, or storing at least one memory snapshot associated with the instruction segment.
10. The method according to claim 1, wherein the method includes storing both the first trace data and the second trace data.
11. A computer system comprising: one or more processors; and one or more computer-readable media having computer-executable instructions stored thereon, the computer-executable instructions, when executed by the one or more processors, cause the computer system to perform at least the following: access trace data representing execution of a plurality of executable instructions of an entity; based on the trace data, identify original information including sensitive information, the original information being accessed based on execution of one or more first executable instructions of the entity; identify derived information of the original information, the identification being based on execution of one or more second executable instructions of the entity that operate on the original information to create the derived information; based on the derived information derived from the original information, determine that the derived information also includes sensitive information; and based on the derived information including the sensitive information, perform one or more of the following: store first trace data including alternative information instead of the derived information, while ensuring that the execution path adopted by the entity based on the derived information will also be adopted during replay of the entity; or store second trace data that causes one or more alternative executable instructions instead of the one or more second executable instructions of the entity to be executed during replay of the entity.
12. The computer system according to claim 11, wherein the computer-executable instructions further cause the computer system to: identify a copy of the original information in the trace data as including sensitive information, the copy of the original information existing at an execution time prior to when the original information first existed in the trace data, and wherein identifying the original information as including sensitive information is based on identifying the copy of the original information in the trace data as including sensitive information.
13. The computer system according to claim 11, wherein the computer-executable instructions further cause the computer system to: A copy of the original information in the traced data that includes sensitive information exists at an execution time after a first occurrence of the original information in the traced data, and wherein identifying that the original information includes sensitive information is based on identifying that a copy of the original information in the traced data includes sensitive information.
14. The computer system according to claim 11, wherein the computer-executable instructions further cause the computer system to: identify a copy of the original information in the traced data that includes sensitive information, and wherein the copy of the original information and the original information are not related by code continuity.
15. The computer system according to claim 11, wherein identifying that the original information includes sensitive information is based on determining that the original information is of a type selected from the list consisting of: a specific data structure, a specific variable, a specific class, a specific field, a specific function, a specific source file, a specific component, a specific module, or executable instructions.
16. The computer system according to claim 11, wherein the original information is identified as sensitive until a type-specific condition associated with the original information has been satisfied.
17. The computer system according to claim 11, wherein the computer system stores the first traced data, and wherein ensuring that the execution path taken by the entity based on the derived information will also be taken during a replay of the original execution of the entity using the first traced data includes one or more of: recording side effects of one or more instructions, recording one or more alternative instructions, or ensuring that the alternative information will produce the same conditional evaluation result as the derived information.
18. The computer system according to claim 11, wherein the computer system stores the second traced data, and wherein storing the second traced data includes one or more of: replacing the instruction segment with one or more instructions that bypass instruction segmentation, replacing the instruction segment with one or more instructions that replicate side effects of the executed instruction segment, or storing at least one memory snapshot associated with the instruction segment.
19. The computer system according to claim 11, wherein the computer system stores both the first traced data and the second traced data.
20. A computer program product, comprising one or more physical hardware storage devices having computer-executable instructions stored thereon, the computer-executable instructions when executed at a processor cause a computer system to perform at least the following: access traced data that represents the execution of a plurality of executable instructions of an entity; based on the traced data, identify original information that includes sensitive information, the original information being accessed based on the execution of one or more first executable instructions of the entity; identify derived information of the original information, the identification being based on the execution of one or more second executable instructions of the entity that operate on the original information to create the derived information; Based on the derived information obtained from the original information, determine that the derived information also includes sensitive information; and Based on the derived information including the sensitive information, perform one or more of the following: Store first trace data including alternative information instead of the derived information, while ensuring that the execution path adopted by the entity based on the derived information will also be adopted during replay of the entity; or Store second trace data that causes one or more alternative executable instructions rather than the one or more second executable instructions of the entity to be executed during replay of the entity.