Deriving profile data for compiler optimization
Patent Information
- Application Number
- CN202180051387.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-31
- Filing Date
- 2021-08-27
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-08-27
Smart Images

Figure CN115968468B_ABST
Abstract
Description
Background Technology
[0001] This invention generally relates to methods, systems, and computer program products for compiler optimization. More specifically, this invention relates to methods, systems, and computer program products for deriving profile data for compiler optimization.
[0002] Computer programs are typically written in programming languages that are easily understood by experienced programmers, and then translated into machine language that the computer processor understands. For many programming languages, a compiler or interpreter translates the computer program into machine language. Typically, a compiler will attempt to translate the entire program file at once and report errors at the end of the process, while an interpreter will attempt to translate the program file line by line and will stop when it encounters an error.
[0003] For programs written in the Java programming language, the source code is typically first translated into an intermediate language called bytecode, which is then translated into machine code. The primary compiler for Java is Javac, which converts Java source code into bytecode organized in class files. A utility called a class loader then loads the bytecode into the Java Virtual Machine (JVM). The JVM includes an interpreter and a Just-In-Time (JIT) compiler, which translates the bytecode and provides the resulting machine code to the computer's processor for running the program. Summary of the Invention
[0004] Illustrative embodiments provide for deriving profile data for compiler optimization. Embodiments include a compiler requesting a first profile dataset associated with a first code segment in response to execution of the first code segment. Embodiments also include performing a query process in response to receiving an indication that the first profile dataset is unavailable, the query process searching for other code segments based on specified criteria relating to attributes of the first code segment. Implementations also include receiving search results from the query process, wherein the search results include a second code segment. Embodiments also include generating an extrapolation profile dataset at least partially based on the second code segment. Embodiments also include storing the extrapolation profile dataset in memory such that the extrapolation profile dataset is associated with the first code segment in memory. Embodiments also include performing an optimization process on the first code segment by a compiler at least partially based on the extrapolation profile dataset. Other embodiments of this aspect include corresponding computer systems, apparatuses, and computer programs recorded on one or more computer storage devices, each computer system and computer storage device configured to perform the actions of the embodiments.
[0005] The embodiments include a computer-usable program product. The computer-usable program product includes a computer-readable storage medium and program instructions stored on the storage medium.
[0006] The embodiments include a computer system. The computer system includes a processor, a computer-readable storage device, and a computer-readable storage medium, as well as program instructions stored on the storage medium for execution by the processor via the memory. Attached Figure Description
[0007] Features considered novel to be part of the invention are set forth in the appended claims. However, the invention itself, as well as its preferred modes of use, further objects and advantages, will be best understood by reading in conjunction with the accompanying drawings and by referring to the following detailed description of illustrative embodiments, wherein: Figure 1 A block diagram of a network that can implement the illustrative embodiments of the data processing system is depicted; Figure 2 A block diagram of a data processing system that can implement illustrative embodiments is depicted; Figure 3 A block diagram of an example computing system according to an illustrative embodiment is depicted; Figure 4 A block diagram depicting an example layout of a virtual machine memory according to an illustrative embodiment; Figure 5 A block diagram of an example virtual machine according to an illustrative embodiment is depicted; Figure 6 A flowchart of an example JIT compiler according to an illustrative embodiment is depicted; Figure 7 A flowchart of an example profile derivation unit according to an illustrative embodiment is depicted; Figure 8 A block diagram of an example virtual machine for deriving an alternative example profile according to an illustrative embodiment is depicted; Figure 9 A flowchart depicts an example process for generating profile data according to an illustrative embodiment; and Figure 10 A flowchart is provided illustrating an alternative example process for generating profile data according to an illustrative embodiment. Detailed Implementation
[0008] Software applications are typically written in source code, which is a high-level representation of the application that is easy for programmers to write and understand. Specialized programs are then used to transform the application into another format that the machine can directly execute. These specialized programs can include compilers, interpreters, and hybrid systems that combine compilers and interpreters. However, for convenience, unless otherwise specified, these specialized programs will be referred to as compilers herein, and similarly, the transformation process will be referred to as compilation herein.
[0009] Generally, compilers typically perform the compilation process either before runtime or during runtime. Each of these options offers certain advantages. For example, runtime compilers allow for portability, automatic memory management, and dynamic loading of code. On the other hand, the processing overhead of runtime compilation has the potential to lead to performance degradation compared to applications compiled before runtime. Therefore, runtime compilers often also use optimization algorithms to mitigate the performance impact of runtime compilation.
[0010] Optimization algorithms used at runtime typically include profile-guided optimizations that perform optimizations based on profile data. Profile data usually contains information summarizing the behavior of instructions at a specific profile point in the executing application. The compiler uses profile data to determine when and where to compile and optimize code segments, such as methods. For example, in some cases, profile data includes a count of how many times a particular instruction has been executed by the virtual machine. This count allows the compiler to estimate the future frequency of execution of that code segment. If the count exceeds a threshold, the compiler applies optimizations to that code segment. Other types of profile data include information about the type of objects that have passed type tests, the target of virtual method dispatches, and the length of arrays and strings.
[0011] Two common techniques for collecting profile data are sampling and programmatic instrumentation. Sampling involves periodically collecting profile data using a timer, while programmatic instrumentation involves using triggers added to the application, such as counter increments. When the compiler queries profile data, it requests such profile data for a specific method and bytecode offset. If profile data is already logged for the specific method and bytecode offset, it is returned in response to the request. Otherwise, if no profile data is available for the specific method and bytecode offset, a sentinel value indicating missing profile data is returned to the compiler in response to the request, and as a result, the compiler cannot perform the expected profile-guided optimizations.
[0012] Thus, the illustrative embodiments recognize that the lack of profile data hinders the compiler's ability to perform profile-guided optimizations. For example, without profile information, the compiler cannot identify which parts of the program code are executed most frequently (e.g., repetitive sequences of execution within the program), which impedes runtime performance. Generally, slow runtime performance is undesirable because it leads to low program throughput, high memory overhead, or other suboptimal behavior. For example, in one test implementation, the x86 machine experienced a 40% reduction in peak throughput without profile information. Therefore, if profile information is unavailable to the compiler, it typically results in a significant performance degradation.
[0013] The illustrative embodiments include those that address this challenge by adapting the compiler optimization process to reduce instances where optimizations are not performed due to a lack of profile data. The illustrative embodiments recognize that alternative information can be used for compiler optimization when profile data is unavailable. For example, in the illustrative embodiments, such alternative information includes program state information, such as information from the state of the class hierarchy table, or extrapolated profile information, such as profile information extrapolated from other parts of the program. The illustrative embodiments recognize that compiler optimizations using such alternative information provide significant performance improvements compared to previous optimizations due to the complete absence of profile information.
[0014] The illustrative embodiments described herein relate to the Java programming language, the Java Virtual Machine (“JVM”), a JIT compiler, an interpreter, and the Java Runtime Environment. However, alternative embodiments may be used in conjunction with other programming languages, virtual machine architectures, or runtime environments. Therefore, for example, terms described in Java terminology, such as “method”, may be used interchangeably with other terms, such as “function”. Furthermore, the term “method” is also synonymous with the terms “class method” and “object method”. A method is a set of code referenced by name and invoked (called) at various points in a program, which prompts the code to execute the method.
[0015] In some embodiments, in the absence of a code segment profile, the compiler uses program state information to optimize the code segment. For example, in some such embodiments, the compiler uses the state of the class hierarchy table to optimize the code segment. An example of this optimization involves calls to potentially polymorphic methods, such as calls to virtual methods. For example, in an abstract class "Shape" with a subclass "Circle" and another subclass "Square", Circle implements a virtual method "draw" that generates a display of a circular shape, while Square implements a virtual method "draw" that generates a display of a square shape. If the array is declared to hold shape objects and is iterated using the "draw" method called on each object, the implementation of each call to the virtual method "draw" will have different results depending on whether the object is a circle or a square. In some embodiments, the compiler determines whether a particular derived class is typically loaded and optimizes the code segment accordingly. For example, if Square is always loaded or most frequently loaded, the compiler may inline the Square implementation and inject "guard code" including type tests to confirm that the object is a Square. If so, the inline method code can be executed; otherwise, an unoptimized virtual method call to another object type is resumed.
[0016] In some such embodiments, the compiler handling calls to such virtual methods accesses data representing the current state of the class hierarchy table to assess whether any of the derived classes have been loaded or instantiated. If the compiler determines that one or more derived classes have been loaded but never instantiated, the compiler only inlines the implementation of the parent class of the virtual method, since none of the derived classes are viable dispatch targets.
[0017] In some such embodiments, the instantiation of derived classes is tracked by profiling at the point where the object is instantiated. In alternative embodiments, the instantiation of derived classes is tracked by processing the application heap at the point where a snapshot is taken. In other alternative embodiments, the instantiation of derived classes is tracked by combining data about which classes are instantiated by which methods (e.g., as tracked in the interpreter) with data from a call graph constructed to capture control flow from the snapshot point. In such embodiments, the compiler combines dynamic execution state (e.g., data about which classes are instantiated by which methods) with static analysis (e.g., data from the snapshot point, such as data from the call graph) to produce an alternative to a lack of profiling information at a certain code location. Having a call graph and performing reachability analysis on the instantiated classes that substantially begin at the code location where the class is instantiated will determine which methods an instance of a given class can reach. This context-sensitive approach, which uses program state at snapshot points, advantageously filters out instantiated classes that cannot reach a particular program point based on reachability analysis.
[0018] In some embodiments, when the compiler examines the profile information of a specific code segment and determines that the profile information is unavailable, the compiler initiates a query process to search for other code segments in the program code that meet specified criteria. If the query process returns code segments that meet the specified criteria, the compiler uses the code segments found by the query process to generate extrapolated profile information.
[0019] In some embodiments, sometimes referred to herein as syntactic matching embodiments, the search criteria include criteria for syntactic matching of a particular code segment. In alternative embodiments, sometimes referred to herein as semantic matching embodiments, the search criteria include criteria for semantic matching of a particular code segment. In other alternative embodiments, sometimes referred to herein as context matching embodiments, the search criteria include criteria for context-sensitive matching of a particular code segment. In still other alternative embodiments, the search criteria include different combinations of syntactic, semantic, and / or context-sensitive matching of a particular code segment.
[0020] In some syntactic matching implementations, the compiler processing a specific code segment for optimization initiates a query process targeting path execution frequency information and / or value profile information. In such implementations, the compiler statically constructs a set of syntactic patterns of interest based on the syntax of sequences within a specific code segment.
[0021] For example, for path execution frequency information, a specific code segment may include the following bytecode sequence:
[0022] The compiler determines the presence of any profile information for the ifeq bytecode by querying bytecode offsets and methods. If no profile information exists, the compiler searches other locations in the program code that test the object as an instance of CustomerName (e.g., other code segments in the program code associated with profile information, such as branch offset information for ifeq fed from the bytecode instance). In some embodiments, the compiler performs this search by generating a syntactic search criterion that includes a statically constructed set of syntactic patterns of interest, such as CustomerName instances followed by ifeq.
[0023] In some embodiments, the compiler assigns a unique identifier to each such sequence. Whenever a class is loaded, the compiler scans its bytecode for matches to that set of patterns and stores the methods and bytecode offsets used for any pattern match into a keyed mapping data structure, where the mapping data structure holds tuples of methods and bytecode offsets matching the corresponding keyed patterns. In some such embodiments, the compiler then looks up the pattern ID if profile information is unavailable and subsequently checks the profile information recorded for other locations matching the same pattern. When patterns are parameterized by value or type, the compiler uses a multi-level mapping of pattern IDs to parameters to methods and tuples representing bytecode offsets at matching locations.
[0024] As another example, for value profile information, a specific code segment may contain the following sequence:
[0025] The compiler, based on the syntactic search of the above code segment, finds other code locations with the same code pattern, calls toString on the declared type java / lang / Object, and uses the merged value profile information to determine whether CustomerRecord.toString should be inlined. If so, the compiler inlines the code at locations without any profile information using parsed guidance.
[0026] In some semantic matching implementations, the compiler processing a specific code segment for optimization statically constructs a set of semantic patterns of interest based on sequences within that specific code segment (e.g., other locations in the program where similar information from different patterns can be provided). For example, a specific code segment may include the following bytecode sequences:
[0027] In some cases, if the compiler attempts syntactic matching, it may not find any other location in the program code where a syntactically matching code segment with profile information exists. Therefore, as an alternative embodiment, the compiler searches for other code segments that can provide similar profile information associated with code segments having different syntactic patterns. For example, the program code may include a code segment associated with profile information, such as:
[0028] In some such embodiments, the compiler stores the method and bytecode offset for any such pattern match into a mapping data structure with a unique identifier of the pattern as the key, where the mapping data structure holds tuples of the method and bytecode offset that match the corresponding keyed pattern. Following this example, in some embodiments, the compiler determines the probability that the receiver is of type CustomerName against another type, and then uses that probability to estimate whether a particular implementation should be inlined as the most likely call target for a toString call.
[0029] In some embodiments, the compiler uses rules for identifying logically related conditions in the code to search for other code segments that can provide similar profile information associated with code segments having different syntactic patterns. For example, a particular code segment may include the following sequence:
[0030] In some such embodiments, the compiler examines the profile information for the code segment described above. If the compiler determines that no profile information exists for that code segment, it then constructs a set of semantic patterns of interest based on the sequence within that particular code segment (e.g., other locations in the program where similar information from different patterns can be provided). The compiler stores the method and bytecode offset for any such pattern matching in a mapping data structure with a unique identifier for the pattern as the key. For example, pattern matching may include the following sequences:
[0031] In some such embodiments, the compiler examines the profile information used for the code segment described above. If the compiler determines that the code segment has profile information, it performs optimizations on the specific code segment being executed. For example, if the profile information indicates a branch deviation, the compiler can use this information to deduce the branch deviation for a specific code segment and optimize that code segment accordingly.
[0032] In some embodiments, the compiler uses rules for seeing through assignments that obscure the fact that two code patterns are actually semantically similar to search for other code segments that can provide similar profile information associated with code segments that have different syntactic patterns. For example, a particular code segment may include the following sequence:
[0033] In some such embodiments, the compiler examines the profile information for the code segment described above. If the compiler determines that no profile information exists for the code segment, it then constructs a set of semantic patterns of interest based on the rules and sequences within that particular code segment. For example, pattern matching may include the following sequences:
[0034] In some such embodiments, the compiler examines the profile information used for the code segment described above. If the compiler determines that the code segment has profile information, it performs optimizations on the specific code segment being executed. For example, if the profile information indicates the type of an object, the compiler can use that information to deduce the type of the object in the specific code segment and optimize the specific code segment by performing profile-guided inlining on the deduced object type.
[0035] In some such embodiments, the compiler uses any of a variety of pattern recognition techniques to find semantically similar code, such as reach definition, abstract interpretation, or value propagation. Such semantic matching embodiments allow for a more complete matching of code segments than syntactic matching by searching based on the semantics of the code segments, with less concern for the actual structure of the code.
[0036] In some context-matching implementations, the compiler processing a specific code segment for optimization will identify candidate code segments, for example, based on syntactic or semantic matching implementations. When the compiler performs a syntactic or semantic search, it is searching for other operations that appear to be doing the same thing as the specific code segment undergoing optimization. For example, the compiler might search the "draw" method of the abstract Shape class to determine the probability that Shape is Square. However, sometimes the problem arises because there can be many different context-relevant and context-unrelevant locations that test for square shapes or give you the probability that Shape is Square. For example, an application might be placed as part of a larger code segment that draws different composite shapes within application code that calls a "draw" method for Square. These locations for square test shapes include the type of information the compiler is looking for. However, some of those locations will test for the same composite shape as the specific code segment undergoing optimization and will therefore be context-similar, while others will test for different composite shapes and will therefore be context-dissimilar. This situation can lead to unpredictable results because the type-test probability can vary significantly depending on which composite shape is being drawn. Therefore, if the compiler simply aggregates similar code syntactically or semantically, the result may be too uncertain to be used as a substitute for profile data.
[0037] Therefore, context-matching embodiments identify candidate code segments, for example, based on syntactic matching or semantic matching embodiments, and then perform context-sensitive analysis of each of the candidate code segments by evaluating the code segments surrounding the candidate code segments. In such embodiments, the compiler further searches for specific important operations preceding and / or following the candidate code segments. In some embodiments, the length of the operations is set or limited according to tunable parameters. The longer the length of the searched operation, the more expensive the processing to search for it for a match, and the less likely the compiler is to find a match, but any search results found are more likely to produce reliable information.
[0038] In some context matching implementations, the compiler generates matching patterns that include syntactic or semantic search criteria for a specific code segment undergoing optimization, and criteria associated with trace information, including operation lengths as prefixes and / or suffixes of the specific code segment. The compiler performs a syntactic or semantic search to identify candidate code segments that match the search criteria and have profiling information. For each candidate code segment, the compiler generates a trace for that segment, which includes the sequence of instructions preceding and / or following the candidate code segment. Once the compiler has generated traces for each candidate code segment, it examines each trace to identify other patterns that match the trace criteria, thereby producing a trace of identified and profiled patterns for each candidate code segment.
[0039] In some context matching embodiments, once the compiler has generated traces of identified and profiled patterns of candidate code segments, it inserts these traces into a data structure for the candidate code segments. In some embodiments, the arrangement of the data structure is based at least in part on the corresponding similarity between a particular code segment and each candidate code segment. For example, in some embodiments, the compiler inserts these traces into a suffix tree data structure to allow for fast search and fuzzy matching. The compiler then begins at points where no profile information is available, pattern matching of the particular code segment, generating traces from the prefix and suffix operations of the particular code segment, and mapping the traces to pattern traces. The compiler takes this pattern trace and uses a known suffix tree matching algorithm to match the pattern trace with one or more existing traces that have profile data. The compiler then uses the profile data of the matched traces to derive alternative profile information for the particular code segment undergoing optimization. For example, the particular code segment may include a sequence of methods that include methods with profile data and methods that do not have profile data:
[0040] In this example, the call to arg2.toString() lacks a profiling information to tell the compiler which implementation of toString is most likely to be called. In some context-matching implementations, the compiler selects traces, such as those including all statements shown in the sample code above. In this example, the compiler has instances of matching operations and patterns of virtual calls. The compiler takes the traced statements and reduces the statement traces to the following pattern traces:
[0041] The compiler uses the pattern trace to match and search for traces with profile data, attempting to find the closest match to the profile data of the second virtual_call in the matching sequence. The compiler can then use the matching trace to provide receiver profile information to be aggregated and used to predict possible dispatch targets for the toString call.
[0042] For clarity of description and without implying any limitation thereof, some example configurations are used to describe illustrative embodiments. Based on this disclosure, those skilled in the art will be able to conceive of many changes, adaptations, and modifications to the described configurations for achieving the described objectives, and the same changes, adaptations, and modifications are contemplated within the scope of the illustrative embodiments.
[0043] Furthermore, simplified diagrams of the data processing environment are used in the accompanying drawings and illustrative embodiments. In a real computing environment, additional structures or components, not shown or described herein, or structures or components that function similarly to those shown but described herein, may exist without departing from the scope of the illustrative embodiments.
[0044] Furthermore, illustrative embodiments are described merely as examples, for specific actual or hypothetical components. For instance, the steps described in different illustrative embodiments can be adapted to provide explanations for decisions made by machine learning classifier models.
[0045] Any particular manifestation of these and other similar products is not intended to limit the invention. Any suitable manifestation of these and other similar products may be chosen within the scope of the illustrative embodiments.
[0046] The examples in this disclosure are for clarity of description only and are not intended to limit the illustrative embodiments. Any advantages listed herein are merely examples and are not intended to limit these illustrative embodiments. Additional or different advantages may be achieved through specific illustrative embodiments. Furthermore, a particular illustrative embodiment may have some, all, or none of the advantages listed above.
[0047] Furthermore, illustrative embodiments can be implemented with respect to any type of data, data source, or access to a data source via a data network. Within the scope of this invention, any type of data storage device can provide data to embodiments of the invention locally within a data processing system or via a data network. Within the scope of the illustrative embodiments, when describing embodiments using a mobile device, any type of data storage device suitable for use with a mobile device can provide data to this embodiment locally on the mobile device or via a data network.
[0048] The exemplary embodiments described using specific code, comparative explanations, computer-readable storage media, advanced features, historical data, designs, architectures, protocols, layouts, diagrams, and tools are merely examples and are not limited to the exemplary embodiments. Furthermore, for clarity, specific software, tools, and data processing environments are used in some instances as examples to describe illustrative embodiments. Illustrative embodiments may be used in conjunction with other comparable or similar structures, systems, applications, or architectures. For example, other similar mobile devices, structures, systems, applications, or architectures may be used in conjunction with such embodiments of the invention within the scope of this invention. Illustrative embodiments may be implemented in hardware, software, or a combination thereof.
[0049] The examples in this disclosure are for clarity of description only and are not intended to limit the scope of the illustrative embodiments. Other data, operations, actions, tasks, activities, and manipulations will arise from this disclosure, and the same data, operations, actions, tasks, activities, and manipulations are contemplated within the scope of the illustrative embodiments.
[0050] Any advantages listed herein are merely examples and are not intended to limit these illustrative embodiments. Additional or different advantages may be achieved through specific illustrative embodiments. Furthermore, a particular illustrative embodiment may have some, all, or none of the advantages listed above.
[0051] Refer to the attached diagram and for details. Figure 1 and 2 These figures are example diagrams of the data processing environment in which illustrative embodiments can be implemented. Figure 1 and 2 This is merely an example and is not intended to assert or imply any limitation regarding the environments in which different embodiments may be implemented. Specific implementations may be modified in many ways based on the environment depicted in the following description.
[0052] Figure 1 A block diagram of a data processing system network in which illustrative embodiments can be implemented is shown. Data processing environment 100 is a computer network in which illustrative embodiments can be implemented. Data processing environment 100 includes network 102. Network 102 is a medium for providing communication links between different devices and computers connected together within data processing environment 100. Network 102 may include connections such as wired, wireless communication links, or fiber optic cables.
[0053] The client or server is merely an example role for certain data processing systems connected to network 102 and is not intended to exclude other configurations or roles of these data processing systems. Data processing system 104 is coupled to network 102. Software applications can execute on any data processing system in data processing environment 100. (Described as...) Figure 1 Any software application executing in processing system 104 can be configured to execute in another data processing system in a similar manner. Figure 1 Any data or information stored or generated in data processing system 104 can be configured to be stored or generated in another data processing system in a similar manner. Data processing systems (such as data processing system 104) may contain data and may have software applications or software tools that perform computational processes thereon. In one embodiment, data processing system 104 includes memory 124, which includes application 105A, which can be configured to implement one or more of the data processor functions described herein according to one or more embodiments.
[0054] Server 106 is coupled to network 102 along with storage unit 108. Storage unit 108 includes database 109 configured to store data as described herein with respect to different embodiments, such as image data and attribute data. Server 106 is a conventional data processing system. In one embodiment, server 106 includes application 105B, which can be configured to implement one or more of the processor functions described herein according to one or more embodiments.
[0055] Clients 110, 112, and 114 are also coupled to network 102. A conventional data processing system, such as server 106 or clients 110, 112, or 114, may contain data and may have software applications or software tools on which conventional computational processes are performed.
[0056] This is merely an example and does not imply any limitations on such an architecture. Figure 1 Certain components available in the example implementations of the embodiments are depicted. For example, server 106 and clients 110, 112, and 114 are depicted as servers and clients only by way of example and are not intended to imply a limitation on the client-server architecture. As another example, the embodiments may be distributed across several data processing systems and the data network shown, while within the scope of the illustrative embodiments, another embodiment may be implemented on a single data processing system. Conventional data processing systems 106, 110, 112, and 114 also represent example nodes in clusters, partitions, and other configurations suitable for implementing the embodiments.
[0057] Device 132 is an example of a conventional computing device described herein. For example, device 132 may take the form of a smartphone, tablet computer, laptop computer, client 110 in fixed or portable form, wearable computing device, or any other suitable device. In one embodiment, device 132 sends a request to server 106 for one or more data processing tasks to be performed by neural network application 105B, such as initiating the neural network's processes as described herein. Figure 1 Any software application running in another conventional data processing system can be configured to run in device 132 in a similar manner. Figure 1 Any data or information stored or generated in another conventional data processing system can be configured to be stored or generated in device 132 in a similar manner.
[0058] Server 106, storage unit 108, data processing system 104, and clients 110, 112, and 114, as well as device 132, can be coupled to network 102 using wired connections, wireless communication protocols, or other suitable data connectivity. Clients 110, 112, and 114 can be, for example, personal computers or network computers.
[0059] In the depicted example, server 106 can provide clients 110, 112, and 114 with data such as boot files, operating system images, and applications. In this example, clients 110, 112, and 114 can be clients of server 106. Clients 110, 112, 114, or some combination thereof, can include their own data, boot files, operating system images, and applications. Data processing environment 100 can include additional servers, clients, and other devices not shown.
[0060] In the depicted example, memory 124 can provide data, such as boot files, operating system images, and applications, to processor 122. Processor 122 may include its own data, boot files, operating system images, and applications. Data processing environment 100 may include additional memory, processor, and other devices not shown.
[0061] In the depicted example, data processing environment 100 can be the Internet. Network 102 can represent a collection of networks and gateways that communicate with each other using Transmission Control Protocol / Internet Protocol (TCP / IP) and other protocols. The core of the Internet is the skeleton of data communication links between master nodes or host computers, including thousands of commercial, government, educational, and other computer systems that route data and messages. Of course, data processing environment 100 can also be implemented as many different types of networks, such as, for example, intranets, local area networks (LANs), or wide area networks (WANs). Figure 1 This is intended as an example, not as an architectural limitation for different illustrative embodiments.
[0062] Among other uses, the data processing environment 100 can be used to implement a client-server environment in which illustrative embodiments can be implemented. The client-server environment enables software applications and data to be distributed across a network, allowing applications to function by interacting with a conventional client data processing system and a conventional server data processing system. The data processing environment 100 can also employ a service-oriented architecture, where interoperable software components distributed across a network can be encapsulated together as a consistent business application. The data processing environment 100 can also take the form of a cloud and employ a service-delivered cloud computing model to enable convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage devices, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or interaction with service providers.
[0063] See Figure 2 The figure depicts a block diagram of a data processing system in which illustrative embodiments can be implemented. The data processing system 200 is an example of a conventional computer, such as... Figure 1 The data processing system 104, server 106, or client 110, 112, and 114, or another type of device in which computer-usable program code or instructions for implementing the illustrative embodiments may reside.
[0064] Data processing system 200 also refers to a conventional data processing system or a configuration thereof, such as Figure 1 The conventional data processing system 132 in which computer-usable program code or instructions for implementing the processing of the illustrative embodiments may be located. The data processing system 200 is described as a computer by way of example only and is not limited thereto. Other devices (e.g., Figure 1 An embodiment in the form of device 132) may modify the data processing system 200, for example, by adding a touch interface, and even eliminate certain depicted components from the data processing system 200 without departing from the general description of the operation and function of the data processing system 200 described herein.
[0065] In the depicted example, the data processing system 200 employs a central architecture including a Northbridge and Memory Controller Hub (NB / MCH) 202 and a Southbridge and Input / Output (I / O) Controller Hub (SB / ICH) 204. A processing unit 206, main memory 208, and a graphics processor 210 are coupled to the Northbridge and Memory Controller Hub (NB / MCH) 202. The processing unit 206 may contain one or more processors and may be implemented using one or more heterogeneous processor systems. The processing unit 206 may be a multi-core processor. In some implementations, the graphics processor 210 may be coupled to the NB / MCH 202 via an Accelerated Graphics Port (AGP).
[0066] In the depicted example, a local area network (LAN) adapter 212 is coupled to the Southbridge and I / O controller hub (SB / ICH) 204. An audio adapter 216, a keyboard and mouse adapter 220, a modem 222, a read-only memory (ROM) 224, a universal serial bus (USB) and other ports 232, and a PCI / PCIe device 234 are coupled to the Southbridge and I / O controller hub 204 via bus 238. A hard disk drive (HDD) or solid-state drive (SSD) 226 and a CD-ROM 230 are coupled to the Southbridge and I / O controller hub 204 via bus 240. The PCI / PCIe device 234 may include, for example, an Ethernet adapter, an insert card, and a PC card for a notebook computer. PCI uses a card bus controller, while PCIe does not. The ROM 224 may be, for example, a flash binary input / output system (BIOS). Hard disk drive 226 and CD-ROM 230 can use, for example, integrated drive electronics (IDE), serial advanced technology accessory (SATA) interface, or variants such as external SATA (eSATA) and micro SATA (mSATA). Super I / O (SIO) device 236 can be coupled to the southbridge and I / O controller hub (SB / ICH) 204 via bus 238.
[0067] Memory such as main memory 208, ROM 224, or flash memory (not shown) are some examples of computer-usable storage devices. Hard disk drives or solid-state drives 226, CD-ROMs 230, and other similar available devices are some examples of computer-usable storage devices that include computer-usable storage media.
[0068] The operating system runs on processing unit 206. The operating system coordinates and provides... Figure 2 The data processing system 200 controls various components within it. The operating system can be a commercially available operating system for any type of computing platform, including but not limited to server systems, personal computers, and mobile devices. Object-oriented or other types of programming systems can operate in conjunction with the operating system and provide calls to the operating system from programs or applications executing on the data processing system 200.
[0069] Operating systems, object-oriented programming systems, and applications or programs (such as...) Figure 1The instructions for application 105 are located on a storage device (such as in the form of code 226A on hard disk drive 226) and can be loaded into at least one of one or more memories (such as main memory 208) for execution by processing unit 206. Processing in an exemplary embodiment can be executed by processing unit 206 using computer-implemented instructions that may reside in memory, such as, for example, main memory 208, read-only memory 224, or one or more peripheral devices.
[0070] Furthermore, in one scenario, code 226A can be downloaded from remote system 201B via network 201A, where similar code 201C is stored on storage device 201D. In another scenario, code 226A can be downloaded to remote system 201B via network 201A, where the downloaded code 201C is stored on storage device 201D.
[0071] Figure 1-2 The hardware within can vary depending on the implementation. (Except for or replacing...) Figure 1-2 The hardware described herein can be replaced with other internal hardware or peripheral devices, such as flash memory, equivalent non-volatile memory, or optical disc drives. Furthermore, the processes described in the exemplary embodiments can be applied to multiprocessor data processing systems.
[0072] In some illustrative examples, the data processing system 200 may be a personal digital assistant (PDA), which is typically configured with flash memory to provide non-volatile memory for storing operating system files and / or user-generated data. The bus system may include one or more buses, such as a system bus, I / O bus, and PCI bus. Of course, the bus system can be implemented using any type of communication structure or architecture that provides data transfer between different components or devices attached to the structure or architecture.
[0073] The communication unit may include one or more devices for sending and receiving data, such as a modem or network adapter. Memory may be, for example, main memory 208 or a cache, such as the cache found in the northbridge and memory controller hub 202. The processing unit may contain one or more processors or CPUs.
[0074] Figure 1-2 The examples depicted and those described above are not intended to imply architectural limitations. For example, the data processing system 200 could take the form of a tablet computer, laptop computer, or telephone device, in addition to being a mobile or wearable device.
[0075] When a computer or data processing system is described as a virtual machine, virtual device, or virtual component, the virtual machine, virtual device, or virtual component operates in a manner similar to data processing system 200, using virtualized representations of some or all of the components depicted in data processing system 200. For example, in a virtual machine, virtual device, or virtual component, processing unit 206 is represented as a virtualized instance of all or some of the hardware processing units 206 available in the host data processing system, main memory 208 is represented as a virtualized instance of all or some of the main memory 208 available in the host data processing system, and disk 226 is represented as a virtualized instance of all or some of the disk 226 available in the host data processing system. In this case, the host data processing system is represented by data processing system 200.
[0076] See Figure 3 This figure depicts a block diagram of an example computing system 300 according to an illustrative embodiment. The figure illustrates an embodiment of the computing system 300 including a virtual machine 314 (such as a JVM). In some embodiments, the virtual machine 314 is configured according to... Figure 9 Or the flowchart shown in 10 is executed. In some embodiments, virtual machine 314 is Figure 1 Examples of applications of 105A / 105B.
[0077] In some embodiments, computing system 300 further includes compiler 304 and operating system 312. Operating system 312 runs on computer hardware and provides an operating environment for applications run by virtual machine 314. Virtual machine 314 includes class loader subsystem 316, interpreter 318, JIT compiler 320, and memory 322. Embodiments of memory 322 may include any form of electronic or computer-usable storage device. In some embodiments, the functionality described herein is distributed across multiple systems, which may include combinations of software- and / or hardware-based systems, such as application-specific integrated circuits (ASICs), computer programs, or smartphone applications.
[0078] In some embodiments, the computing system 300 includes source code 302, which contains code written in a specific programming language, such as Java, C, C++, C#, Ruby, or Perl. Thus, the source code 302 conforms to a specific set of syntactic and / or semantic rules for the associated language. For example, code written in Java conforms to the Java language specification. However, because specifications are updated and modified over time, the source code 302 may be associated with a version number indicating the version of the specification that the source code 302 follows.
[0079] In some embodiments, compiler 304 transforms source code 302 into class files 306, such as bytecode, representing a program to be executed. For example, in some embodiments, source code 302 includes one or more files input to compiler 304. Compiler 304 processes one or more files and outputs the resulting bytecode as one or more class files 306.
[0080] In some embodiments, class file 306 is then loaded and executed by execution engine 308, which includes runtime environment 310 and operating system 312. Runtime environment 310 includes virtual machine 314, which receives class file 306 and inputs the bytecode of class file 306 into interpreter 318 or JIT compiler 320. Interpreter 318 and JIT compiler 320 generate machine code that virtual machine 314 outputs to operating system 312 for running the program. In some embodiments, virtual machine 314 processes the bytecode of the program while the program is executing. Therefore, the processing time of virtual machine 314 affects the runtime performance of the program. In some embodiments, JIT compiler 320 uses optimization techniques described herein to reduce the processing time of virtual machine 314, thereby improving the runtime performance of the program.
[0081] In some embodiments, virtual machine 314 uses interpreter 318 and JIT compiler 320 to execute a program using a combination of interpretation and compilation techniques. In some such embodiments, virtual machine 314 initially begins by using interpreter 318 to interpret the bytecode representing the program while collecting profile data related to program behavior. For example, some embodiments collect profile data representing the frequency with which different parts or blocks of code are executed by virtual machine 314. In some embodiments, when an interpreted method call is executed, virtual machine 314 increments a "call counter," which is metadata stored in a profile structure associated with the called method. When the call counter exceeds a compilation threshold, virtual machine 314 spawns a compilation thread (or uses an existing compilation thread) to compile / optimize the method. In an alternative embodiment, virtual machine 314 tracks the number of loop iterations in a method using a "tail-edge counter" and triggers compilation when the tail-edge counter reaches a compilation threshold.
[0082] In some embodiments, once a code block exceeds a specified threshold (is “hot”), the virtual machine 314 invokes the JIT compiler 320 to use optimization techniques to reduce the amount of processing time that the virtual machine 314 uses to process the code block into machine-level instructions. However, some such embodiments include an unoptimized program execution period during which the virtual machine 314 processes the code block (e.g., by executing the code in an interpreted mode) until JIT compilation before the block is identified as a hot code block.
[0083] In some embodiments, the virtual machine 314 implements the techniques described herein to perform compiler optimizations during program execution, including when little or no profile data is available for one or more code segments or blocks. Such program execution periods are not limited to the early or initial stages of program execution but can occur throughout the program's execution. For example, due to the powerful ability to dynamically load classes, Java-based programs can change in real-time at any point during program execution. This dynamic loading of classes invalidates compiler optimizations and profile data collected by the virtual machine 314 prior to class changes. When compiler optimizations and profile data become invalid due to dynamic loading, this creates a program execution cycle that can occur at any point during program execution, during which little or no profile data is available for the code segments associated with the dynamically loaded class. In some embodiments, the virtual machine 314 uses the techniques described herein, which allow compiler optimizations of code segments during any such program execution period. Such embodiments have the potential to improve program performance throughout program execution compared to embodiments lacking optimizations during such periods.
[0084] While the illustrated embodiment shows memory 322 as part of virtual machine 314, alternative embodiments may locate memory 322 elsewhere outside virtual machine 314, but in communication with virtual machine 314. Embodiments of memory 322 may include memory on a single electronic or computer-usable storage device or memory distributed among any number of electronic or computer-usable storage devices.
[0085] See Figure 4 This figure depicts a block diagram of an exemplary layout of virtual machine memory 400 according to an illustrative embodiment. While components of virtual machine memory 400 may be illustrated or referred to as memory regions, blocks, or otherwise, it is not required that the memory regions be contiguous. In some embodiments, virtual machine memory 400 is... Figure 3 Example of memory 322.
[0086] In the illustrated embodiment, virtual machine memory 400 includes a heap 402, a class data area 404, a thread data area 412, and a method area 418. In this embodiment, heap 402 represents a runtime data area from which memory for class instances and arrays is allocated. In this embodiment, class data area 404 represents a memory area storing data about each individual class. In this embodiment, for each loaded class, class data area 404 includes a runtime constant pool 406 representing data from a constant table for each class, method code 408 representing virtual machine instructions for methods of the class, and field and method data 410 representing data such as static field data for each class. Method area 418 includes class data shared by all threads running across the virtual machine, including a runtime constant pool 420.
[0087] In the illustrated embodiment, thread data region 412 represents a memory region where structures specific to each thread are stored. For virtual machines (e.g., ...), Figure 3 Each thread executing on the virtual machine (314) has a thread region 412 including a thread stack 414 and a program counter 416. In one embodiment, the program counter 416 stores the current address of the virtual machine instruction being executed for each corresponding thread. Therefore, as a thread steps through an instruction, the program counter is updated to maintain the index of the current instruction. In one embodiment, the thread stack 414 stores frames for its corresponding thread, which hold local variables and partial results, and are also used for method calls.
[0088] See Figure 5 This figure depicts a block diagram of an example virtual machine 500 according to an illustrative embodiment. In some embodiments, virtual machine 500 is Figure 3 Example of virtual machine 314.
[0089] In the illustrated embodiment, the virtual machine 500 includes a JIT compiler 502 and an interpreter 504. While the illustrated embodiment depicts the JIT compiler 502 and interpreter 504 as separate components, in alternative embodiments, the JIT compiler 502 and interpreter 504 may be integrated into a single component. In some embodiments, the virtual machine 500 includes a decision mechanism 506 that controls whether bytecode will be compiled by the JIT compiler 502 or interpreted by the interpreter 504. In some embodiments, the decision mechanism 506 selects the JIT compiler 502 based on the availability of profile data 524 for the called segment. However, the embodiments disclosed herein allow for the derivation of profile information for code segments lacking profile data. Such embodiments allow the decision mechanism to select the JIT compiler 502 even when the profile data 524 is unavailable or insufficient for a given called code segment. For example, in some embodiments, the decision mechanism 506 selects the JIT compiler 502 based on the type of code in the called code segment, regardless of the availability of the profile data memory 524.
[0090] In the illustrated embodiment, the virtual machine 500 includes a class loader subsystem 514, which provides mechanisms for loading types, namely classes and interfaces, and makes them accessible to the JIT compiler 502 and the interpreter 504. In the illustrated embodiment, the virtual machine 500 also includes a frequency data component 520 and a type data component 522, which are examples of profiling instruments for obtaining information for constructing profile data in a profile data memory 524. The type data component 522 collects type information about each variable in the program and records the type information in the profile data memory 524. The frequency data component 520 collects frequency information indicating how many times each function is executed and records this information in the profile data memory 524.
[0091] In some embodiments, JIT compiler 502 performs a second-level compilation to create code (e.g., JIT code) as optimized code output and executed by execution component 516, which updates data heap 518 and executes the application on operating system 526. In some embodiments, JIT compiler 502 uses profile data from profile data storage 524 to optimize code segments. In some embodiments, if no profile data is available for a given code segment, JIT compiler 502 will instead attempt to optimize the code segment using alternative data, such as class hierarchy data 512 or data from application heap 510 at the point where application snapshot 508 was taken (e.g., the snapshot is preserved in the application's state).
[0092] In some embodiments, there are fail-safe paths that couple execution component 516 to each of JIT compiler 502 and interpreter 504. These fail-safe paths pass control from execution component 516 to interpreter 504, which does not make such assumptions, when certain assumptions in the optimized code (e.g., optimized JIT code) become invalid during runtime and execution needs to fall into the interpreter's hands.
[0093] See Figure 6 This figure depicts a block diagram of an example JIT compiler 600 according to an illustrative embodiment. While the illustrated embodiment shows a JIT compiler 600, alternative embodiments include other types of optimizing compilers. In some embodiments, the JIT compiler 600 is Figure 5 An example of the JIT compiler 502.
[0094] In some embodiments, the JIT compiler 600 includes a detector 602 that receives code segment 612, for example, as bytecode, and detects or derives information about the code segment, such as the type of instructions represented by the bytecode. For example, the bytecode may relate to calls to potentially polymorphic methods, such as calls to virtual methods, or other types of code with the potential to be optimized by the JIT compiler 600. In the illustrated embodiment, the detector 602 sends information about the code segment 612 detected by the detector 602 to a profile request module 604. The profile request module 604 receives the information about the code segment 612, which triggers the profile request module 604 to request profile data for the code segment 612. For example, in some embodiments, the profile request module 604 requests profile information by sending a request for profile information associated with a particular method and a bytecode offset associated with the code segment 612. In some embodiments, the functionality described herein is distributed across multiple systems, which may include a combination of software- and / or hardware-based systems, such as application-specific integrated circuits (ASICs), computer programs, or smartphone applications.
[0095] In some embodiments, the JIT compiler 600 includes a decision module 606 that determines whether the profile request module 604 has received any profile information in response to a request. If so, the profile data is provided to the optimizing compiler 608. Otherwise, the profile derivation unit 610 is triggered by the lack of profile data to generate data that can be used as an alternative for optimization in the absence of profile data. For example, in some embodiments, the profile derivation unit 610 receives information about code segment 612 detected by detector 602 and associated state information 614 as the basis for generating alternative profile data. In some embodiments, the profile derivation unit 610 receives the information detected by detector 602 directly from detector 602, while in alternative embodiments, the profile derivation unit 610 receives the information detected by detector 602 forwarded from the profile request module 604. The alternative profile data is then provided to the optimizing compiler 608. The optimizer compiler 608 receives profile data or profile alternative data and uses any data it receives to create optimized executable code, which is then output by the optimizer compiler 608 to the execution component (e.g., ...). Figure 5 The execution component 516) is provided for use in the operating system (e.g., Figure 5 It runs on the operating system 526.
[0096] See Figure 7 The figure depicts a block diagram of an example profile derivation unit 700 according to an illustrative embodiment. In some embodiments, the profile derivation unit 700 is Figure 6 Example of the simplified derivation unit 610.
[0097] In the illustrated embodiment, the profile derivation unit 700 receives code segment information 706 associated with the code segment being compiled. For example, in one embodiment, the code segment information 706 is information received by the profile derivation unit 610 from the detector 602 regarding... Figure 6 An instance of information detected by detector 602 in code segment 612.
[0098] In the illustrated embodiment, the profile derivation unit 700 uses state information 708 to optimize the code segment in the absence of profile information for the code segment. For example, in some such embodiments, the compiler uses class hierarchy data (e.g., Figure 5 We can optimize code segments by using the class hierarchy data (512).
[0099] In the illustrated embodiment, the profile derivation unit 700 includes a query builder 702, a query engine 704, and a result processor 710. In some embodiments, the functionality described herein is distributed across multiple systems, which may include a combination of software- and / or hardware-based systems, such as application-specific integrated circuits (ASICs), computer programs, or smartphone applications.
[0100] In some embodiments, query builder 702 evaluates code segment information 706 to generate criteria for search state information 708. For example, in some embodiments, query builder 702 generates one or more criteria for searching data representing the current state of a class hierarchy table. Query engine 704 receives one or more criteria from query builder 702 and initiates searches against one or more data sources. Result processor 710 receives the results of one or more queries initiated by query engine 704 and filters the results, making the data usable for the optimization process. Result processor 710 provides the filtered results as optimization data 714 to the optimization compiler, such as... Figure 6 The optimizing compiler 608. In some embodiments, the result processor 710 generates optimized data 714, the format of which is... Figure 6 The profile request module 604 outputs the same profile data.
[0101] As an example, in one embodiment, query builder 702 identifies code segment information 706 involving calls to potentially polymorphic methods, such as calls to virtual methods, and thus generates a query for state information 708 to determine whether a particular derived class is typically loaded so that the code segment can be optimized accordingly. In some such embodiments, query builder 702 generates any derived classes that have been loaded as a criterion for the query. In alternative embodiments, query builder 702 also generates any derived classes that have been instantiated as another criterion for the query. In some embodiments, query engine 704 receives one or more criteria (searching for loaded derived classes) from query builder 702 and initiates a search for state information 708 (such as class hierarchy data). Result processor 710 receives the results of one or more queries initiated by query engine 704, which in this example may include a list of loaded classes, and alternatively may also include a list of instantiated classes. Result processor 710 filters these results for data that can be used for the optimization process, for example by determining that there exists a particular derived class that is always or most frequently called from loaded and / or instantiated classes. The result processor 710 provides the filtered results as optimization data 714 to the optimization compiler, such as... Figure 6 The optimizing compiler 608, in this example, can inline the normally called derived class and inject "guide code" including type tests to confirm that the object is a derived class. In some embodiments, the result processor 710 produces optimized data 714, whose format is similar to... Figure 6 The profile request module 604 outputs the same profile data.
[0102] In some such embodiments, at the point of object instantiation, the profiling instrument 722 (such as...) is used... Figure 5 The instantiation of derived classes is tracked using frequency data components 520. In an alternative embodiment, the instantiation of derived classes is tracked by processing the application heap 718 at the point where snapshot 716 is taken. In other alternative embodiments, the instantiation of derived classes is tracked by combining data about which classes are instantiated by which methods (e.g., as tracked in interpreter 724) with data from call graph 726 or other class hierarchy data 720 constructed to capture control flow from snapshot 716. In such embodiments, the compiler combines dynamic execution state (e.g., data about which classes are instantiated by which methods) with static analysis (e.g., data from snapshot points, such as data from call graph 726) to produce an alternative to a lack of profiling information at a certain code location. A reachability analysis of classes instantiated at the point where call graph 726 is taken and substantially at the code location where the class is instantiated reveals which methods an instance of a given class can reach. This context-sensitive approach, which uses program state at snapshot points, advantageously filters out instantiated classes that cannot reach a particular program point based on reachability analysis.
[0103] refer to Figure 8 The figure depicts a block diagram of an alternative example profile derivation unit 800 according to an illustrative embodiment. In some embodiments, the profile derivation unit 800 is Figure 6 Example of the simplified derivation unit 610.
[0104] In the illustrated embodiment, the profile derivation unit 800 receives code segment information 820 associated with the code segment being compiled. For example, in one embodiment, the code segment information 820 is information received by the profile derivation unit 610 from the detector 602 regarding... Figure 6 An instance of information detected by detector 602 in code segment 612.
[0105] In the illustrated embodiment, the profile derivation unit 800 uses state information 822 to optimize the code segment in the absence of profile information for the code segment. For example, in some such embodiments, the compiler uses class hierarchy data (e.g., Figure 5 We can optimize code segments by using the class hierarchy data (512).
[0106] In the illustrated embodiment, the profile derivation unit 800 includes a query builder 802, a query engine 804, a query result processor 806, a trace generator 808, a trace analyzer 810, a context matching module 812, a context result processor 814, and a memory 818. In some embodiments, the functionality described herein is distributed across multiple systems, which may include a combination of software- and / or hardware-based systems, such as application-specific integrated circuits (ASICs), computer programs, or smartphone applications.
[0107] In some embodiments, query builder 802 evaluates code segment information 820 to generate criteria for searching state information 822. For example, in some embodiments, query builder 802 generates one or more criteria for searching data representing the current state of a class hierarchy table. Query engine 804 receives one or more criteria from query builder 802 and initiates a search for state information 822 and / or one or more other data sources. Query results processor 806 receives the results of one or more queries initiated by query engine 804 and stores the obtained code segments as candidate code segments 836 in memory 818. Query results processor 806 also provides state information 822 to trace generator 808.
[0108] Trace generator 808 retrieves statement traces, including context segments, from application code before and after the invoked code segment. Trace generator 808 then stores the statement traces as trace data 838 in memory 818. In some embodiments, trace generator 808 then retrieves the corresponding result context trace for each of the candidate code segments 836 from the query results, wherein the result context trace includes result context segments from application code before and after the result code segment. In some embodiments, trace analyzer 810 reduces the context segments to the invoked pattern trace. In some embodiments, trace analyzer 810 then generates a suffix tree data structure from the candidate code segments 836 and their corresponding pattern traces. Trace analyzer 810 then stores the suffix tree data structure as suffix tree data 840 in memory 818. Context matching module 812 maps the invoked pattern traces to the suffix tree to identify the closest matching results and result context segments with relevant profile data. The context matching module 812 provides the closest match result and the resulting context code segment to the context result processor 814. The context result processor 814 generates an extrapolation profile dataset based on the data from the closest match result found by the context matching module 812. The context result processor 814 then provides this data to the optimizing compiler (such as...). Figure 6 The optimizing compiler 608 provides an extrapolated profile dataset. In this example, the optimizing compiler can inline normally invoked derived classes and inject "guide code" including type tests to confirm that the object is a derived class. In some embodiments, the context result processor 814 produces results consistent with those generated by... Figure 6 The profile request module 604 outputs profile data with the same optimized data 842 format.
[0109] In some such embodiments, at the point of object instantiation, by means of, Figure 5The instantiation of derived classes is tracked by the profiling instrument 830 of the frequency data component 520. In an alternative embodiment, the instantiation of derived classes is tracked by processing the application heap 826 at the point where snapshot 824 is taken. In other alternative embodiments, the instantiation of derived classes is tracked by combining data about which classes are instantiated by which methods (e.g., as tracked in interpreter 832) with data from call graph 834 or by constructing additional class hierarchy data 828 to capture control flow from snapshot 824. In such embodiments, the compiler combines dynamic execution state (e.g., data about which classes are instantiated by which methods) with static analysis (e.g., data from snapshot points, such as data from call graph 834) to produce an alternative to missing profiling information at a certain code location. With call graph 834, and substantially performing reachability analysis on instantiated classes starting at the code location where the instantiated class is instantiated, it reveals which methods an instance of a given class can reach. This context-sensitive approach, which uses program state at snapshot points, advantageously filters out instantiated classes that cannot reach a particular program point based on reachability analysis.
[0110] See Figure 9 The figure depicts a flowchart of an example process 900 for generating profile data according to an illustrative embodiment. In a specific embodiment, this is a compiler execution process 900, such as a JIT compiler 502.
[0111] In this embodiment, at box 902, the compiler receives the code segment that triggered the compilation call. Next, at box 904, the compiler requests profile data for the called code segment. Next, at box 906, the compiler receives a response indicating that the profile data is unavailable for the called code segment. Next, at box 908, the compiler performs a query procedure that searches for the parsed code segment based on specified criteria associated with the attributes of the called code segment. Next, at box 910, the compiler receives search results from the query procedure, including the resulting code segment. Next, at box 912, the compiler generates an extrapolated profile dataset based on the relevant profile data. Next, at box 914, the compiler stores the extrapolated profile dataset in memory such that the extrapolated profile dataset is associated with the called code segment in memory. Next, at box 916, the compiler performs an optimization procedure on the called code segment based on the extrapolated profile dataset.
[0112] See Figure 10 The figure depicts a flowchart of an example process 1000 for generating profile data according to an illustrative embodiment. In a specific embodiment, this is a compiler execution process 1000, such as a JIT compiler 502.
[0113] In one embodiment, at box 1002, the compiler receives the invoked code segment that triggered compilation. Next, at box 1004, the compiler requests profile data for the invoked code segment. Next, at box 1006, the compiler receives a response indicating that the profile data is unavailable for the invoked code segment. Next, at box 1008, the compiler retrieves statement traces, including the context code segment, from the application code preceding and following the invoked code segment. Next, at box 1010, the compiler narrows down the context code segment to the invoked pattern trace. Next, at box 1012, the compiler performs a query procedure that searches for the parsed code segment based on specified criteria related to the attributes of the invoked code segment. Next, at box 1014, the compiler receives search results, including the result code segment, from the query procedure. Next, at box 1016, the compiler retrieves the result context code segment from the application code preceding and following the result code segment. Next, at box 1018, the compiler generates a suffix tree data structure from candidate code segment 836 and the result context code segments preceding and following the result code segment. Next, at box 1020, the compiler maps the called pattern trace to the suffix tree to identify the closest matching result and the result context code segment with relevant profile data. Next, at box 1022, the compiler generates an extrapolation profile dataset based on the relevant profile data. Next, at box 1024, the compiler stores the extrapolation profile dataset in memory, associating it with the called code segment in memory. Next, at box 1026, the compiler performs an optimization process on the called code segment based on the extrapolation profile dataset.
[0114] The following definitions and abbreviations will be used to interpret the claims and description. As used herein, the terms “comprising,” “including,” “having,” “containing,” “or any other variation thereof” are intended to cover non-exclusive inclusion. For example, a composition, mixture, process, method, article, or apparatus that comprises a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or other elements inherent to such composition, mixture, process, method, article, or apparatus.
[0115] Furthermore, the term "illustrative" is used herein to mean "serving as an example, illustration, or illustration." Any embodiment or design described herein as "illustrative" is not necessarily to be construed as preferred or advantageous over other embodiments or designs. The terms "at least one" and "one or more" should be understood to include any integer greater than or equal to one, i.e., one, two, three, four, etc. The term "multiple" should be understood to include any integer greater than or equal to two, i.e., two, three, four, five, etc. The term "connection" can include both indirect "connection" and direct "connection."
[0116] References to "an embodiment," "embodiment," "exemplary embodiment," etc., in this specification indicate that the described embodiment may include a particular feature, structure, or characteristic; however, each embodiment may or may not include a particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same implementation. Moreover, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is believed that the influence of other embodiments (whether explicitly described or not) on such feature, structure, or characteristic is within the knowledge of those skilled in the art.
[0117] The terms “about,” “substantially,” “roughly,” and their variations are intended to include the degree of error associated with a measurement of a specific quantity based on equipment available at the time of application submission. For example, “about” could include a range of ±8%, 5%, or 2% of a given value.
[0118] Various embodiments of the invention have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or technical improvements over those found in the market, or to enable those skilled in the art to understand the embodiments described herein.
[0119] Various embodiments of the invention have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or technical improvements over those found in the market, or to enable those skilled in the art to understand the embodiments described herein.
[0120] Therefore, computer-implemented methods, systems, or apparatuses, as well as computer program products, for managing participation and other related features, functions, or operations in online communities are provided in the illustrative embodiments. When embodiments or portions thereof are described with respect to the type of apparatus, the computer-implemented methods, systems, or apparatuses, computer program products, or portions thereof are adapted or configured for use with suitable and comparable performance to that type of apparatus.
[0121] Where embodiments are described as being implemented within an application, the delivery concept of an application in a Software as a Service (SaaS) model is contemplated within the scope of the illustrative embodiments. In a SaaS model, the capabilities of an application implementing an embodiment are provided to a user by executing the application within a cloud infrastructure. Users can access the application using various client devices through thin client interfaces such as web browsers (e.g., web-based email) or other lightweight client-applications. Users do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, or storage of the cloud infrastructure. In some cases, users may not even manage or control the capabilities of the SaaS application. In some other cases, the SaaS implementation of the application may allow for limited user-specific application configuration settings that may be anomalous.
[0122] This invention can be a system, method, and / or computer program product with any possible level of technical detail integration. The computer program product may include one or more computer-readable storage media having computer-readable program instructions thereon for causing a processor to execute aspects of the invention.
[0123] Computer-readable storage media can be tangible means for retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital universal disk (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or protrusions in slots having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.
[0124] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.
[0125] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions by utilizing state information from the computer-readable program instructions to personalize the electronic circuitry in order to perform aspects of this invention.
[0126] The present invention will be described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0127] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium storing the instructions includes an article of manufacture containing instructions that implement aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0128] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce computer-implemented processing, such that the instructions executed on the computer, other programmable apparatus, or other device perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than indicated in the figures. For example, depending on the functions involved, two consecutively shown blocks may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0130] Various embodiments of the present invention can also be delivered as part of service engagement with client companies, non-profit organizations, government entities, internal organizational structures, etc. Aspects of these embodiments may include configuring computer systems to perform and deploy some or all of the software, hardware, and web services implementing the methods described herein. Aspects of these embodiments may also include analyzing client operations, creating recommendations in response to the analysis, building a system to implement portions of the recommendations, integrating the system into existing processes and infrastructure, metering system usage, allocating expenses to users of the system, and accounting for system usage. For example, a computer program product for metering the system may further include program instructions for metering the use of the program instructions associated with the request; and program instructions for generating invoices based on the metered usage. While the above embodiments of the invention have been described by setting forth their respective advantages, the invention is not limited to their specific combinations. Rather, such embodiments can be combined in any manner and number according to the intended deployment of the invention without loss of their beneficial effects.
Claims
1. A computer-implemented method for deriving abbreviated data, comprising: The compiler requests the first profile dataset associated with the first code segment in response to the execution of the first code segment; In response to receiving an indication that the first profile dataset is unavailable, a query process is executed, which searches for other code segments based on specified criteria related to the attributes of the first code segment; Receive search results from the query process, wherein the search results include a second code segment; At least in part, an extrapolation profile dataset is generated based on the second code segment; The extrapolation profile dataset is stored in memory such that the extrapolation profile dataset is associated with the first code segment in the memory; as well as The compiler performs an optimization process on the first code segment based at least in part on the extrapolation profile dataset.
2. The computer-implemented method according to claim 1, wherein, The query process also searches for program status information.
3. The computer-implemented method according to claim 2, wherein, The program status information includes a class hierarchy table.
4. The computer-implemented method according to claim 1, wherein, The specified standard includes a syntactic standard that is at least partially based on the syntax of the first code segment.
5. The computer-implemented method according to claim 4, wherein, The specified criteria include syntactic matching of the first code segment.
6. The computer-implemented method according to claim 4, wherein, The specified criteria include rules describing the semantic matching of the first code segment.
7. The computer-implemented method according to claim 1, wherein, The query process includes: The candidate code segment is identified based on semantic matching between the candidate code segment and the first code segment; Retrieve the context code segment from the application code before and after the first code segment; Retrieve the resulting context code segment from the application code preceding and following the candidate code segment; and Based on the fuzzy matching between the result context code segment and the context code segment, the candidate code segment is identified as the second code segment.
8. The computer-implemented method according to claim 7, wherein, The query process includes generating a data structure for candidate code segments, wherein the arrangement of the data structure is based at least in part on the corresponding similarity between the first code segment and each of the candidate code segments.
9. The computer-implemented method according to claim 1, wherein, The query process includes limiting the search results to code segments that have associated profile datasets.
10. The computer-implemented method according to claim 9, wherein, Generating the extrapolation profile dataset includes generating the extrapolation profile dataset of the first code segment based at least in part on the second profile dataset of the second code segment.
11. A computer program product for deriving profile data, the computer program product comprising program instructions executable by a processor to cause the processor to perform operations including: The compiler requests the first profile dataset associated with the first code segment in response to the execution of the first code segment; In response to receiving an indication that the first profile dataset is unavailable, a query process is executed, which searches for other code segments based on specified criteria related to the attributes of the first code segment; Receive search results from the query process, wherein the search results include a second code segment; At least in part, an extrapolation profile dataset is generated based on the second code segment; The extrapolation profile dataset is stored in memory such that the extrapolation profile dataset is associated with the first code segment in the memory; as well as The compiler performs an optimization process on the first code segment based at least in part on the extrapolation profile dataset.
12. The computer program product according to claim 11, wherein, The program instructions are stored in a computer-readable storage device in the data processing system, and the stored program instructions are transmitted from a remote data processing system via a network.
13. The computer program product according to claim 11, wherein, The program instructions are stored in a computer-readable storage device in a server data processing system, and wherein the stored program instructions are downloaded via a network to a remote data processing system for use in a computer-readable storage device associated with the remote data processing system, the computer program product further comprising: Program instructions for measuring the use of the program instructions associated with the request; and Program instructions for generating invoices based on measured usage.
14. The computer program product according to claim 11, wherein, The query process also searches for program status information, which includes a class hierarchy table.
15. The computer program product according to claim 11, wherein, The specified criteria include a syntactic criterion at least in part based on the syntax of the first code segment, and wherein the specified criteria include syntactic matching of the first code segment.
16. The computer program product according to claim 11, wherein, The query process includes: The candidate code segment is identified based on semantic matching between the candidate code segment and the first code segment; Retrieve the context code segment from the application code before and after the first code segment; Retrieve the resulting context code segment from the application code preceding and following the candidate code segment; and Based on the fuzzy matching between the result context code segment and the context code segment, the candidate code segment is identified as the second code segment.
17. A computer system for deriving profile data, comprising a processor and one or more computer-readable storage media, wherein program instructions are stored on the one or more computer-readable storage media, the program instructions being executable by the processor to cause the processor to perform operations including: The compiler requests the first profile dataset associated with the first code segment in response to the execution of the first code segment; In response to receiving an indication that the first profile dataset is unavailable, a query process is executed, which searches for other code segments based on specified criteria related to the attributes of the first code segment; Receive search results from the query process, wherein the search results include a second code segment; At least in part, an extrapolation profile dataset is generated based on the second code segment; The extrapolation profile dataset is stored in memory such that the extrapolation profile dataset is associated with the first code segment in the memory; as well as The compiler performs an optimization process on the first code segment based at least in part on the extrapolation profile dataset.
18. The computer system according to claim 17, wherein, The query process also searches for program status information, including a class hierarchy table.
19. The computer system according to claim 17, wherein, The specified criteria include a syntactic criterion at least in part based on the syntax of the first code segment, and wherein the specified criteria include syntactic matching of the first code segment.
20. The computer system according to claim 17, wherein, The query process includes: The candidate code segment is identified based on semantic matching between the candidate code segment and the first code segment; Retrieve the context code segment from the application code before and after the first code segment; Retrieve the resulting context code segment from the application code preceding and following the candidate code segment; and Based on the fuzzy matching between the result context code segment and the context code segment, the candidate code segment is identified as the second code segment.
Citation Information
Patent Citations
Profile guided jit code generation
CN103064720A
Method, system, and computer program product for extending sparse partial redundancy elimination to support speculative code motion within an optimizing compiler
US6151706A