Lazy compression in garbage collection
Through lazy compression technology, using lazy free lists and loading barriers, only relocating surviving objects when necessary, solving the performance waste problem caused by relocating the imminent dead objects in the garbage collection system, and improving system efficiency and cache locality.
Patent Information
- Application Number
- CN202380081610.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-17
- Filing Date
- 2023-10-13
- Publication Date
- 2025-07-08
AI Technical Summary
When compressing memory, existing garbage collection systems need to relocate a large number of objects that are about to die, resulting in waste of performance and inefficiency.
Using lazy compression methods, through lazy free list and loading barrier technology, relocate surviving objects only when necessary, reducing unnecessary relocation operations.
Improves garbage collection performance, reduces unnecessary relocation overhead, improves the overall efficiency and cache locality of the system, and reduces the latency of application threads.
Smart Images

Figure CN120283228A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] The following applications are incorporated herein by reference: Application No. 17 / 967,214, filed on October 17, 2022; Application No. 17 / 303,634, filed on June 3, 2021; Application No. 17 / 303,635, filed on June 3, 2021; Application No. 17 / 303,636, filed on June 3, 2021. Technical Field
[0003] This disclosure relates to garbage collection in computer systems. Specifically, this disclosure relates to compaction during garbage collection. Background Art
[0004] A compiler converts source code written according to a specification for the convenience of a programmer into machine code (also referred to as "native code" or "object code"). The machine code can be directly executed by a physical machine environment. Additionally or alternatively, the compiler converts the source code into an intermediate representation (also referred to as "virtual machine code / instructions") (such as bytecode), which can be executed by a virtual machine capable of running on various physical machine environments. The virtual machine instructions can be executed by the virtual machine in a more direct and efficient manner than the source code. Converting the source code into virtual machine instructions includes mapping the source code functionality according to a specification to virtual machine functionality that utilizes the underlying resources (such as data structures) of the virtual machine. Generally, a functionality presented by a programmer in simple terms via the source code is converted into more complex steps that more directly map to the instruction set supported by the underlying hardware on which the virtual machine resides.
[0005] A virtual machine executes applications and / or programs by executing an intermediate representation (such as bytecode) of the source code. The interpreter of the virtual machine converts the intermediate representation into machine code. When an application is executed, certain memory (also referred to as "heap memory") is allocated to the objects created by the program. A garbage collection system can be used to automatically reclaim the memory locations occupied by objects that the application no longer uses. The garbage collection system frees the programmer from having to explicitly specify which objects to deallocate. The generational garbage collection scheme is based on the empirical observation that most objects are only used for short periods of time. In generational garbage collection, two or more allocation regions (generations) are assigned and kept separate based on the age of the objects contained therein. New objects are created in the "young" generation that is recycled regularly; and when a generation is full, the objects still referenced by one or more objects stored in the older generation regions are copied to (i.e., "promoted to") the next oldest generation. Sometimes a full scan is performed.
[0006] Memory used by a garbage collector can become fragmented over time; that is, free memory may be scattered between the locations of the remaining live objects. Fragmentation of free memory means that new objects may not be allocated contiguously. To reduce fragmentation, the garbage collector can compact the memory by relocating the remaining live objects, leaving contiguous free memory in which to allocate new objects. Relocating objects can easily consume 40% or more of the garbage collector's time. However, many of the objects relocated by the garbage collector (in some cases, 90% or more) will die before the next garbage collection cycle, meaning that the time spent relocating those objects is effectively wasted. If the garbage collector does not relocate as many objects that are about to die during compaction, then the performance of the garbage collector can be significantly improved.
[0007] The methods described in this section are methods that can be practiced, but not necessarily methods that have been previously contemplated or practiced. Thus, unless otherwise indicated, no assumption should be made that any of the methods described in this section are prior art merely by virtue of their inclusion in this section. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Embodiments are illustrated in the figures of the accompanying drawings by way of example and not by way of limitation. References to "one" or "an" embodiment in this disclosure do not necessarily refer to the same embodiment and mean at least one. In the accompanying drawings:
[0009] Figure 1 An example computing architecture in which the techniques described herein may be practiced is illustrated.
[0010] Figure 2 is a block diagram illustrating one embodiment of a computer system suitable for implementing the methods and features described herein.
[0011] Figure 3 An example virtual machine memory layout in block diagram form according to an embodiment is illustrated.
[0012] Figure 4 An example frame in block diagram form according to an embodiment is illustrated.
[0013] Figure 5 The execution engine and heap memory of a virtual machine according to an embodiment are illustrated.
[0014] Figure 6 Heap references and dereferencable references according to an embodiment are illustrated.
[0015] Figure 7 A reference load barrier according to an embodiment is illustrated.
[0016] Figure 8 Illustrates a reference write barrier according to an embodiment.
[0017] Figure 9 Illustrates a set of example operations for lazy compression according to one or more embodiments;
[0018] Figure 10 Illustrates an example of lazy compression according to one or more embodiments; and
[0019] Figure 11 Shows a block diagram illustrating a computer system according to one or more embodiments. Detailed Description
[0020] In the following description, for purposes of explanation and to provide a thorough understanding, numerous specific details are set forth. One or more embodiments may be practiced without these specific details. Features described in one embodiment may be combined with features described in different embodiments. In some examples, well-known structures and devices are described in block diagram form to avoid unnecessarily obscuring the present invention.
[0021] The following table of contents is provided for reference purposes only and should not be construed as limiting the scope of one or more embodiments.
[0022] 1. General Overview
[0023] 2. Architecture Overview
[0024] 2.1. Example Architecture
[0025] 2.2. Example Class File Structure
[0026] 2.3. Example Virtual Machine Architecture
[0027] 2.4. Loading, Linking, and Initialization
[0028] 3. Garbage Collection
[0029] 4. Loading and Write Barriers
[0030] 5. Lazy Compression
[0031] 6. Example Embodiments
[0032] 7. Computer Networks and Cloud Networks
[0033] 8. Hardware Overview
[0034] 9. Miscellaneous; Extensions
[0035] 1. General Overview
[0036] One or more embodiments eliminate a significant amount of the overhead associated with relocating objects during garbage collection by performing “lazy” compaction. When the garbage collector selects memory regions to be scanned in the next relocation phase, the garbage collector adds those regions to a “lazy free list” (LFL). When the allocator needs to allocate memory, the allocator preferentially uses the list of confirmed free regions. However, if the confirmed free list is exhausted, the allocator can select regions from the LFL and perform the (one or more) load barriers associated with any objects marked as live in the destination region. If an object is still live when the object's load barrier is executed, the load barrier relocates the object. Thus, performing the (one or more) load barriers associated with any objects marked as live in the destination region results in a fully contiguous memory region for allocation.
[0037] Lazy compaction can improve system performance in various ways. For example:
[0038] 1. Concurrent garbage collection (GC) algorithms typically reserve some memory “head room” to avoid running out of memory before the GC has time to finish. In cases where the head room can accommodate the most expensive regions for relocation marking, the GC can avoid incurring the cost of most of the relocation work.
[0039] 2. Since the load barriers are triggered when loading an object that needs to be relocated, lazy compaction relocates objects in access order to a “to-space” (described below), which can help improve the cache locality of the object graph.
[0040] 3. When the GC relocates objects in a subsequent GC cycle, the relocated portions of the object graph can be laid out in an approximate depth-first search order, which can also help improve the cache locality of the object graph (compared to maintaining the allocation order in a compressed form).
[0041] 4. As the application threads get closer to exhausting the available memory, some GCs force the application threads to sleep to give the GC time to “catch up”. However, this sleep time can cause an unacceptable degradation in application performance (e.g., to the extent that the application fails to meet a service level agreement (SLA)). As described herein, lazy compaction introduces a small amount of overhead per allocation, which may degrade performance just enough for the GC to keep up without causing significant latency.
[0042] 5. When the GC performs relocation during the relocation phase, the GC must redundantly bring the live set into the cache again, only into the cache and then not use the live set again. When an application thread performs relocation during lazy compaction, the application thread brings it into a cache that is about to be used by allocation anyway, thus reducing cache coherence traffic.
[0043] 6. In the GC method that performs eager relocation, a forwarding table data structure that contains information about where objects are relocated can be fully allocated from the start. With lazy compaction, the data structure can be dynamically allocated and grow as needed. Since only a small fraction of all objects will ultimately be relocated during the relocation phase, the memory wasted due to an unnecessarily large forwarding table can be significantly reduced.
[0044] One or more embodiments described in this specification and / or recited in the claims may not be included in this general overview section.
[0045] 2. Architecture Overview
[0046] 2.1. Example Architecture
[0047] Figure 1 An example architecture is illustrated in which the techniques described herein may be practiced. Software and / or hardware components described in relation to the example architecture may be omitted or associated with different sets of functions than those described herein. Software and / or hardware components not described herein may be used within an environment according to one or more embodiments. Thus, the example environment should not be construed as limiting the scope of any claims.
[0048] As Figure 1 illustrated, computing architecture 100 includes source code file 101, which is compiled by compiler 102 into class file 103 representing the program to be executed. Class file 103 is then loaded and executed by execution platform 112, which includes runtime environment 113, operating system 111, and one or more application programming interfaces (APIs) 110 that enable communication between runtime environment 113 and operating system 111. Runtime environment 113 includes virtual machine 104, which includes various components such as memory manager 105 (which may include a garbage collector), class file validator 106 for checking the validity of class file 103, class loader 107 for locating and constructing in-memory representations of classes, interpreter 108 for executing the code of virtual machine 104, and just-in-time (JIT) compiler 109 for generating optimized machine-level code.
[0049] In an embodiment, computing architecture 100 includes a source code file 101 that contains code written in a particular programming language such as Java, C, C++, C#, Ruby, Perl, etc. Thus, the source code file 101 adheres to a set of specific syntax and / or semantic rules for the associated language. For example, code written in Java adheres to the Java language specification. However, since the specifications are updated and revised over time, the source code file 101 may be associated with a version number that indicates the revision of the specification that the source code file 101 adheres to. The exact programming language used to write the source code file 101 is generally not critical.
[0050] In various embodiments, a compiler 102 converts the source code, written according to the specifications for the convenience of the programmer, into machine or object code that can be directly executed by a particular machine environment, or into an intermediate representation ("virtual machine code / instructions") such as bytecode that can be executed by a virtual machine 104, which is capable of running on top of various particular machine environments. The virtual machine instructions can be executed by the virtual machine 104 in a more direct and efficient manner than the source code. Converting the source code into virtual machine instructions involves mapping the source code functionality from the language to the virtual machine functionality that utilizes underlying resources such as data structures. Generally, the functionality presented by the programmer in simple terms via the source code is converted into more complex steps that more directly map to the instruction set supported by the underlying hardware on which the virtual machine 104 resides.
[0051] Generally speaking, a program is executed as a compiled program or an interpreted program. When a program is compiled, the code is globally transformed from a first language into a second language before execution. Since the work of transforming the code is performed ahead of time, the compiled code tends to have excellent runtime performance. Additionally, since the transformation occurs globally before execution, the code can be analyzed and optimized using techniques such as constant folding, dead code elimination, inlining, etc. However, depending on the program being executed, the startup time may be significant. Additionally, inserting new code will require the program to be taken offline, recompiled, and re-executed. For many dynamic languages (such as Java) that are designed to allow code to be inserted during the execution of the program, a purely compiled approach may not be appropriate. When a program is interpreted, the code of the program is read line by line and converted into machine-level instructions while the program is being executed. Thus, the program has a short startup time (it can start executing almost immediately), but the runtime performance is diminished by performing the on-the-fly transformation. Additionally, since each instruction is analyzed individually, many optimizations that rely on a more global analysis of the program cannot be performed.
[0052] In some embodiments, virtual machine 104 includes an interpreter 108 and a JIT compiler 109 (or components that implement aspects of both) and uses a combination of interpreted techniques and compiled techniques to execute programs. For example, virtual machine 104 may initially begin by interpreting virtual machine instructions that represent a program via interpreter 108 while tracking statistics related to program behavior, such as how often different parts or blocks of code are executed by virtual machine 104. After a code block exceeds a threshold (is "hot"), virtual machine 104 invokes JIT compiler 109 to perform an analysis of the block and generate optimized machine-level instructions that are used to replace the "hot" code block for future execution. Since programs tend to spend most of their time executing a small portion of the overall code, compiling only the "hot" parts of the program can provide performance similar to fully compiled code but without the startup penalty. Additionally, although the optimization analysis is limited to the "hot" block being replaced, there is still far greater optimization potential than transforming each instruction individually. There are many variations of the above example, such as tiered compilation.
[0053] To provide a clear example, source code file 101 has been illustrated as the "top-level" representation of a program to be executed by execution platform 112. Although computing architecture 100 depicts source code file 101 as the "top-level" program representation, in other embodiments, source code file 101 may be an intermediate representation received via a "higher-level" compiler that processes code files in different languages into the language of source code file 101. Some examples in the following disclosure assume that source code file 101 follows a class-based object-oriented programming language. However, this is not a requirement for leveraging the features described herein.
[0054] In an embodiment, compiler 102 receives source code file 101 as input and converts source code file 101 into class file 103 in a format expected by virtual machine 104. For example, in the context of the JVM, the Java Virtual Machine Specification defines a specific class file format that class file 103 is expected to follow. In some embodiments, class file 103 contains virtual machine instructions that have been converted from source code file 101. However, in other embodiments, class file 103 may also contain other structures, such as tables identifying constant values and / or metadata related to various structures (classes, fields, methods, etc.).
[0055] The following discussion assumes that each class file in class file 103 represents a corresponding "class" defined in source code file 101 (or dynamically generated by compiler 102 / virtual machine 104). However, the foregoing assumption is not a strict requirement and will depend on the implementation of virtual machine 104. Accordingly, the techniques described herein can still be performed regardless of the exact format of class file 103. In some embodiments, class file 103 is divided into one or more "libraries" or "packages", each of which includes a collection of classes that provide related functionality. For example, a library may contain one or more class files that implement input / output (I / O) operations, mathematical tools, encryption techniques, graphics utilities, and the like. Additionally, some classes (or fields / methods within those classes) may include access constraints that limit the use of these classes to within a particular class / library / package or to classes with appropriate permissions.
[0056] 2.2. Example Class File Structure
[0057] Figure 2 Illustrated is an example structure for class file 200 in the form of a block diagram according to an embodiment. For the sake of providing a clear example, the remainder of this disclosure assumes that the class files 103 of computing architecture 100 follow the structure of example class file 200 described in this section. However, in a real-world environment, the structure of class file 200 will depend on the implementation of virtual machine 104. Additionally, one or more of the features discussed herein may modify the structure of class file 200 to, for example, add additional structure types. Accordingly, the exact structure of class file 200 is not critical to the techniques described herein. For the purposes of section 2.1, "class" or "current class" refers to the class represented by class file 200.
[0058] In Figure 2 class file 200 includes a constant table 201, a field structure 208, class metadata 207, and a method structure 209. In an embodiment, constant table 201 is a data structure that serves as a symbol table for the class among other functions. For example, constant table 201 may store data related to various identifiers used in source code file 101, such as types, scopes, contents, and / or locations. Constant table 201 has entries for value structures 202 (representing constant values of types such as int, long, double, float, byte, string, etc.) derived from source code file 101 by compiler 102, class information structures 203, name and type information structures 204, field reference structures 205, and method reference structures 206. In an embodiment, constant table 201 is implemented as an array that maps index i to structure j. However, the exact implementation of constant table 201 is not critical.
[0059] In some embodiments, the entries of the constant table 201 include a structure that indexes other entries of the constant table 201. For example, an entry of a value structure 202 representing a string in the value structure 202 can hold a tag that identifies the "type" of the entry as a string and an index to one or more other value structures 202 of the constant table 201, where the one or more value structures store char, byte, or int values representing the ASCII characters of the string.
[0060] In an embodiment, the field reference structure 205 of the constant table 201 holds an index to a class information structure in the constant table 201 that represents the class defining the field in the class information structure 203, and an index to a name and class information structure in the constant table 201 that provides the name and descriptor of the field in the name and type information structure 204. The method reference structure 206 of the constant table 201 holds an index to a class information structure in the constant table 201 that represents the class defining the method in the class information structure 203, and an index to a name and class information structure in the constant table 201 that provides the name and descriptor of the method in the name and type information structure 204. The class information structure 203 holds an index to a value structure in the constant table 201 that holds the name of the associated class in the value structure 202.
[0061] The name and type information structure 204 holds an index to a value structure in the constant table 201 that stores the name of the field / method in the value structure 202, and an index to a value structure in the constant table 201 that stores the descriptor in the value structure 202.
[0062] In an embodiment, the class metadata 207 includes metadata for the class, such as the (one or more) version number, the number of entries in the constant pool, the number of fields, the number of methods, access flags (whether the class is public, private, final, or abstract, etc.), an index to a class information in the class information structure 203 of the constant table 201 that identifies the current class, an index to the class information structure 203 of the constant table 201 that identifies the superclass (if any), etc.
[0063] In an embodiment, the field structure 208 represents a set of structures that identify the various fields of a class. The field structure 208 stores, for each field of the class, an accessor flag for the field (whether the field is static, public, private, or final, etc.), an index to a value structure in the constant table 201 that holds the field name in the value structure 202, and an index to a value structure in the constant table 201 that holds the field descriptor in the value structure 202.
[0064] In an embodiment, the method structure 209 represents a set of structures for various methods that identify classes. The method structure 209 stores, for each method of the class, an accessor flag for the method (whether the method is static, public, private, or synchronized, etc.), an index to a value structure in the constant table 201 that holds the method name in the value structure 202, an index to a value structure in the constant table 201 that holds the method descriptor in the value structure 202, and virtual machine instructions corresponding to the method body as defined in the source code file 101.
[0065] In an embodiment, a descriptor represents the type of a field or method. For example, a descriptor can be implemented as a string that follows a specific syntax. Although the exact syntax is not critical, some examples are described below.
[0066] In an example where the descriptor represents a field type, the descriptor identifies the data type held by the field. In an embodiment, a field can hold a primitive type, an object, or an array. When the field holds a primitive type, the descriptor is a string that identifies the primitive type (e.g., "B" = byte, "C" = char, "D" = double, "F" = float, "I" = int, "J" = long int, etc.). When the field holds an object, the descriptor is a string that identifies the object class name (e.g., "L ClassName"). In this case, "L" indicates a reference, so "L ClassName" represents a reference to an object of the ClassName class. When the field is an array, the descriptor identifies the type held by the array. For example, "[B" indicates a byte array, where "[" indicates an array and "B" indicates that the array holds the primitive type of bytes. However, since arrays can be nested, the descriptor for an array can also indicate nesting. For example, "[[L ClassName" indicates an array where each index holds an array that holds an object of the ClassName class. In some embodiments, ClassName is fully qualified and includes the simple name of the class as well as the path name of the class. For example, ClassName can indicate the location where the file is stored in the package, library, or file system that hosts the class file 200.
[0067] In the case of a method, the descriptor identifies the parameters of the method and the return type of the method. For example, a method descriptor may follow the general form "({ParameterDescriptor})ReturnDescriptor", where {ParameterDescriptor} is a list of field descriptors representing the parameters, and ReturnDescriptor is a field descriptor identifying the return type. For example, the string "V" may be used to represent a void return type. Thus, a method defined in source code file 101 as "Object m(intI, doubled, Thread t) {...}" matches the descriptor "(ID LThread)LObject".
[0068] In an embodiment, the virtual machine instructions held in method structure 209 include operations that reference entries of constant table 201. Using Java as an example, consider the following class:
[0069]
[0070] In the above example, the Java method add12and13 is defined in class A, does not accept any parameters, and returns an integer. The body of the method add12and13 calls the static method addTwo of class B, which accepts constant integer values 12 and 13 as parameters and returns a result. Therefore, in the constant table 201, the compiler 102 includes, among other entries, a method reference structure corresponding to the call to the method B.addTwo. In Java, calls to methods are compiled down to invoke commands in the JVM bytecode (invokestatic in this case, because addTwo is a static method of class B). The invoke command is provided with an index into the constant table 201, which corresponds to a method reference structure that identifies the class "B" that defines addTwo, the name of addTwo "addTwo", and the descriptor "(II)I" of addTwo. For example, assuming that the aforementioned method reference is stored at index 4, the bytecode instruction can appear as "invokestatic#4".
[0071] Since the constant pool 201 refers to classes, methods, and fields of structures carrying identification information in symbolic form, rather than making direct references to memory locations, the entries in the constant pool 201 are referred to as "symbolic references". One reason symbolic references are used in the class file 103 is that, in some embodiments, the compiler 102 does not know how and where these classes will be stored after the classes are loaded into the runtime environment 113. As will be described in Section 2.3, after the referenced classes (and associated structures) have been loaded into the runtime environment and assigned specific memory locations, the runtime representation of the symbolic references is ultimately resolved by the virtual machine 104 into actual memory addresses.
[0072] 2.3. Example Virtual Machine Architecture
[0073] Figure 3 FIG. illustrates an example virtual machine memory layout 300 in block diagram form according to an embodiment. For the sake of providing a clear example, the remaining discussion will assume that the virtual machine 104 follows Figure 3 the virtual machine memory layout 300 depicted in. Additionally, although the components of the virtual machine memory layout 300 may be referred to as memory "areas", there is no requirement that these memory areas be contiguous.
[0074] In Figure 3 the example illustrated, the virtual machine memory layout 300 is divided into a shared area 301 and a thread area 307. The shared area 301 represents the area of memory that stores structures shared among the various threads executing on the virtual machine 104. The shared area 301 includes a heap 302 and a per-class area 303. In an embodiment, the heap 302 represents the runtime data area from which memory for class instances and arrays is allocated. In an embodiment, the per-class area 303 represents the memory area that stores data related to individual classes. In an embodiment, for each loaded class, the per-class area 303 includes a runtime constant pool 304 representing data from the constant pool 201 of the class, field and method data 306 (e.g., to hold the static fields of the class), and method code 305 representing the virtual machine instructions for the methods of the class.
[0075] The thread area 307 represents the memory area that stores structures specific to individual threads. In Figure 3 the, the thread area 307 includes a thread structure 308 and a thread structure 311 representing the per-thread structures utilized by different threads. For the sake of providing a clear example, Figure 3 the thread area 307 depicted in assumes that two threads are executing on the virtual machine 104. However, in a real-world environment, the virtual machine 104 can execute any arbitrary number of threads, and the number of thread structures scales accordingly.
[0076] In an embodiment, the thread structure 308 includes a program counter 309 and a virtual machine stack 310. Similarly, the thread structure 311 includes a program counter 312 and a virtual machine stack 313. In an embodiment, the program counter 309 and the program counter 312 store the current addresses of the virtual machine instructions being executed by their respective threads.
[0077] Thus, as the thread steps through the instructions, the program counter is updated to maintain an index to the current instruction. In an embodiment, the virtual machine stack 310 and the virtual machine stack 313 each store frames for their respective threads, which hold local variables and partial results, and are also used for method calls and returns.
[0078] In an embodiment, a frame is a data structure for storing data and partial results, the return value for a method, and performing dynamic linking. A new frame is created each time a method is called. When the method for which the frame was generated completes, the frame is destroyed. Thus, when a thread executes a method call, the virtual machine 104 creates a new frame and pushes the frame onto the virtual machine stack associated with that thread.
[0079] When the method call completes, the virtual machine 104 passes the result of the method call back to the previous frame and pops the current frame from the stack. In an embodiment, for a given thread, one frame is active at any point. This active frame is called the current frame, the method that caused the current frame to be generated is called the current method, and the class to which the current method belongs is called the current class.
[0080] Figure 4 An example frame 400 in block diagram form according to an embodiment is illustrated. For the sake of providing a clear example, the remaining discussion will assume that the frames of the virtual machine stack 310 and the virtual machine stack 313 follow the structure of the frame 400.
[0081] In an embodiment, the frame 400 includes local variables 401, an operand stack 402, and a runtime constant pool reference table 403. In an embodiment, the local variables 401 are represented as an array of variables, each variable holding a value (such as a boolean, byte, char, short, int, float, or reference). Additionally, some value types (such as long or double) may be represented by more than one entry in the array. The local variables 401 are used to pass parameters on method calls and store partial results. For example, when a frame 400 is generated in response to a method call, the parameters may be stored in predefined locations within the local variables 401, such as indices 1 - N corresponding to the first through Nth parameters in the call.
[0082] In an embodiment, when frame 400 is created by virtual machine 104, the operand stack 402 is initially empty. Then, virtual machine 104 supplies instructions from the method code 305 of the current method to load constants or values from local variables 401 onto the operand stack 402. Other instructions obtain operands from the operand stack 402, operate on those operands, and push the results back onto the operand stack 402. Additionally, the operand stack 402 is used to prepare the arguments to be passed to a method and to receive the method result. For example, the arguments of the method being called can be pushed onto the operand stack 402 before the call to the method is issued. Virtual machine 104 then generates a new frame for the method call, where the operands on the operand stack 402 of the previous frame are popped and loaded into the local variables 401 of the new frame. When the called method terminates, the new frame is popped from the virtual machine stack and the return value is pushed onto the operand stack 402 of the previous frame.
[0083] In an embodiment, the runtime constant pool reference table 403 contains a reference to the runtime constant pool 304 of the current class. The runtime constant pool reference table 403 is used to support resolution. Resolution is the process by which symbolic references in the constant pool 304 are converted into actual memory addresses, classes are loaded as needed to resolve undefined symbols, and variable accesses are translated into appropriate offsets into the storage structures associated with the runtime locations of those variables.
[0084] 2.4. Loading, Linking, and Initialization
[0085] In an embodiment, virtual machine 104 dynamically loads, links, and initializes classes. Loading is the process of finding a class with a specific name and creating a representation in memory within the runtime environment 113 from the associated class file 200 of that class. For example, a runtime constant pool 304, method code 305, and field and method data 306 are created for the class within each-class area 303 of the virtual machine memory layout 300. Linking is the process of obtaining the in-memory representation of the class and combining that representation with the runtime state of virtual machine 104 so that the methods of the class can be executed. Initialization is the process of executing the class constructor to set the initial state of the field and method data 306 and / or to create a class instance for the initialized class on the heap 302.
[0086] The following are examples of loading, linking, and initialization techniques that can be implemented by virtual machine 104. However, in many embodiments, these steps can be interleaved such that an initial class is loaded and then a second class is loaded during linking to resolve symbolic references found in the first class, which in turn causes a third class to be loaded, etc. Thus, the progress through the loading, linking, and initialization phases can vary by class. Additionally, some embodiments can defer ("lazy execute") one or more functions of the loading, linking, and initialization process until a class is actually needed. For example, the resolution of method references can be deferred until the virtual machine instruction to invoke the method is executed. Thus, the exact timing of when these steps are performed for each class can vary widely between different implementations.
[0087] To begin the loading process, virtual machine 104 starts by calling class loader 107 to load the initial class. The technique by which the initial class is specified will vary by embodiment. For example, one technique can cause virtual machine 104 to accept a command line argument specifying the initial class at startup.
[0088] To load a class, class loader 107 parses the class file 200 corresponding to the class and determines whether the class file 200 is well - formed (meets the syntactic expectations of virtual machine 104). If not, class loader 107 generates an error. For example, in Java, the error might be generated in the form of an exception that is thrown to an exception handler for processing. Otherwise, class loader 107 generates an in - memory representation of the class by allocating a runtime constant pool 304, method code 305, and field and method data 306 for the class within each class area 303.
[0089] In some embodiments, when class loader 107 loads a class, class loader 107 also recursively loads the parent class of the loaded class. For example, virtual machine 104 can ensure that the parent class of a particular class has been loaded, linked, and / or initialized before proceeding with the loading, linking, and initialization process for that particular class.
[0090] During linking, virtual machine 104 verifies the class, prepares the class, and performs the resolution of symbolic references defined in the runtime constant pool 304 of the class.
[0091] To verify a class, virtual machine 104 checks whether the in-memory representation of the class is structurally correct. For example, virtual machine 104 can check that every class other than the generic class Object has a superclass, that a final class has no subclasses and a final method is not overridden, that the constant pool entries are consistent with each other, that the current class has the correct access permissions for the classes / fields / methods referenced in the constant pool 304, that the virtual machine 104 code for a method will not result in unexpected behavior (e.g., ensuring that a jump instruction does not send virtual machine 104 out of the end of the method), etc. The exact checks performed during verification depend on the implementation of virtual machine 104. In some cases, verification may cause additional classes to be loaded, but it is not necessarily required to also link those classes before proceeding. For example, assume class A contains a reference to a static field of class B. During verification, virtual machine 104 can check class B to ensure that the referenced static field actually exists, which may cause class B to be loaded, but does not necessarily cause class B to be linked or initialized. However, in some embodiments, certain verification checks can be deferred until a later stage, such as being checked during the resolution of symbolic references. For example, some embodiments can defer checking access permissions for symbolic references until those references are being resolved.
[0092] To prepare a class, virtual machine 104 initializes the static fields located within the field and method data 306 of that class to default values. In some cases, setting the static fields to default values may not be the same as running the class's constructor. For example, the verification process can zero out the static fields, or set the static fields to the values that the constructor would expect those fields to have during initialization.
[0093] During resolution, virtual machine 104 dynamically determines the actual memory addresses from the symbolic references included in the class's runtime constant pool 304. To resolve a symbolic reference, virtual machine 104 uses class loader 107 to load the class identified in the symbolic reference (if not already loaded). After being loaded, virtual machine 104 has knowledge of the memory locations within the per-class area 303 of the referenced class and its fields / methods. Then, virtual machine 104 replaces the symbolic reference with a reference to the actual memory location of the referenced class, field, or method. In an embodiment, virtual machine 104 caches the resolution to be reused in case virtual machine 104 encounters the same class / name / descriptor when processing another class. For example, in some cases, class A and class B can call the same method of class C. Thus, when resolution is performed for class A, the result can be cached and reused during the resolution of the same symbolic reference in class B to reduce overhead.
[0094] In some embodiments, the step of resolving symbolic references during linking is optional. For example, an embodiment may perform symbolic resolution in a "lazy" manner, deferring the step of resolution until the virtual machine instructions for the referenced class / method / field are executed.
[0095] During initialization, the virtual machine 104 executes the class constructor to set the starting state of the class. For example, initialization may initialize the fields and method data 306 of the class and generate / initialize any class instances created by the constructor on the heap 302. For example, the class file 200 of the class may specify that a particular method is the constructor for setting the starting state. Thus, during initialization, the virtual machine 104 executes the instructions of the constructor.
[0096] In some embodiments, the virtual machine 104 performs resolution of field and method references by initially checking whether the field / method is defined in the referenced class. Otherwise, the virtual machine 104 recursively searches the superclasses of the referenced class for the referenced field / method until the field / method location is determined or the top-level superclass is reached, in which case an error is generated.
[0097] 3. Garbage Collection
[0098] Figure 5 Illustrated is an execution engine and a heap memory of a virtual machine according to an embodiment. As Figure 5 illustrated, the system 500 includes an execution engine 502 and a heap 530. The system 500 may include more or fewer components than Figure 5 illustrated components. Figure 5 The components illustrated in
[0099] In one or more embodiments, the heap 530 represents a runtime data area from which memory for class instances and arrays is allocated. An example of the heap 530 was described above as Figure 3 the heap 302 in
[0100] The heap 530 stores objects 534a-d created during the execution of the application. The objects stored in the heap 530 may be ordinary objects, object arrays, or another type of object. An ordinary object is a class instance. Class instances are explicitly created by class instance creation expressions. An object array is a container object that holds a fixed number of values of a single type. An object array is a particular collection of ordinary objects.
[0101] The heap 530 stores live objects 534b, 534d (indicated by the spotted pattern) and unused objects 534a, 534c (also referred to as "dead objects", represented by the blank pattern). Unused objects are objects that are no longer used by any application. Live objects are objects that are still being used by at least one application. An object is still being used by an application if the object (a) is pointed to by a root reference, or (b) is traceable from another object pointed to by a root reference. An object is "traceable" from a second object if a reference to the first object is included in the second object.
[0102] Sample code may include the following:
[0103]
[0104] The application thread 508a that executes the above sample code creates an object temp in the heap 530. The object temp is of type Person and includes two fields. Since the field age is an integer, the portion of the heap 530 allocated to temp directly stores the value "3" of the field age. Since the field name is a string, the portion of the heap 530 allocated to temp does not directly store the value of the name field; rather, the portion of the heap 530 allocated to temp stores a reference to another object of type String. The String object stores the value "Sean". The String object is said to be "traceable" from the Person object.
[0105] In one or more embodiments, the execution engine 502 includes one or more threads configured to perform various operations. For example, as shown, the execution engine 502 includes garbage collection (GC) threads 506a-b and application threads 508a-b.
[0106] In one or more embodiments, the application threads 508a-b are configured to perform operations of one or more applications. The application threads 508a-b create objects that are stored on the heap 530 during runtime. The application threads 508a-b may also be referred to as "mutators" because the application threads 508a-b can modify the heap 530 (during the concurrent phase of the GC cycle and / or between GC cycles).
[0107] In one or more embodiments, the GC threads 506a-b are configured to perform garbage collection. The GC threads 506a-b can iteratively execute GC cycles based on scheduling and / or event triggers (such as when a threshold allocation of the heap (or a region thereof) is reached). A GC cycle includes a set of GC operations for reclaiming memory locations occupied by unused objects in the heap.
[0108] In an embodiment, multiple GC threads 504a-b may perform GC operations in parallel. Multiple GC threads 506a-b that work in parallel may be referred to as a "parallel collector".
[0109] In an embodiment, GC threads 506a-b may perform at least some GC operations concurrently with the execution of application threads 508a-b. GC threads 504a-b that operate concurrently with application threads 508a-b may be referred to as a "concurrent collector" or a "partial concurrent collector".
[0110] In an embodiment, GC threads 506a-b may perform generational garbage collection. The heap is separated into different regions. A first region (which may be referred to as the "young generation space") stores objects that have not met the criteria for being promoted from the first region to a second region; a second region (which may be referred to as the "old generation space") stores objects that have met the criteria for being promoted from the first region to the second region. For example, when a live object has survived at least a threshold number of GC cycles, the live object is promoted from the young generation space to the old generation space.
[0111] Various different GC processes for performing garbage collection achieve different memory efficiencies, time efficiencies, and / or resource efficiencies. In an embodiment, different GC processes may be performed for different heap regions. For example, the heap may include a young generation space and an old generation space. One type of GC process may be performed for the young generation space. A different type of GC process may be performed for the old generation space. Examples of different GC processes are described below.
[0112] As a first example, a copying collector involves at least two separately defined heap address spaces referred to as a "from-space" and a "to-space". The copying collector identifies live objects stored in the area defined as the from-space. The copying collector copies the live objects to another area defined as the to-space. After all live objects have been identified and copied, the area defined as the from-space is reclaimed. New memory allocations may begin at a first location in the original from-space.
[0113] The copying may be done using at least three different regions within the heap: an Eden space and two survivor spaces S1 and S2. Objects are initially allocated in the Eden space. When the Eden space is full, a GC cycle is triggered. Live objects are copied from the Eden space to one of the survivor spaces, such as S1. In the next GC cycle, the live objects in the Eden space are copied to the other survivor space, which would be S2. Additionally, the live objects in S1 are also copied to S2.
[0114] As another example, a mark-sweep garbage collector separates the GC operation into at least two phases: a mark phase and a sweep phase. During the mark phase, the mark-sweep garbage collector marks each live object with a "live" bit. The live bit can be, for example, a bit within the object header of the live object. During the sweep phase, the mark-sweep garbage collector traverses the heap to identify all unmarked blocks of contiguous memory address space. The mark-sweep garbage collector links the unmarked blocks together into an organized free list. The unmarked blocks are reclaimed. New memory allocations are performed using the free list. New objects can be stored in the memory blocks identified from the free list.
[0115] The mark-sweep garbage collector can be implemented as a parallel garbage collector. Additionally or alternatively, the mark-sweep garbage collector can also be implemented as a concurrent garbage collector. Example phases within the GC cycle of a concurrent mark-sweep garbage collector include:
[0116] · Phase 1: Identify objects referenced by root references (this is not concurrent with the executing application)
[0117] · Phase 2: Mark objects reachable from the objects referenced by root references (this can be concurrent)
[0118] · Phase 3: Identify objects that have been modified as part of the execution of the program during Phase 2 (this can be concurrent)
[0119] · Phase 4: Re-mark the objects identified at Phase 3 (this is not concurrent)
[0120] · Phase 5: Sweep the heap to obtain a free list and reclaim memory (this can be concurrent)
[0121] As another example, a compacting garbage collector attempts to compact the reclaimed memory regions. The heap is partitioned into a set of equally sized heap regions, each being a contiguous range of virtual memory. The compacting garbage collector performs a concurrent global marking phase to determine the liveness of objects throughout the heap. After the marking phase is complete, the compacting garbage collector identifies regions that are mostly empty. The compacting garbage collector first reclaims these regions, which typically results in a large amount of free space. The compacting garbage collector focuses its reclaiming and compacting activities on regions of the heap that are likely to be full of reclaimable objects (i.e., garbage). The compacting garbage collector copies the live objects from one or more regions of the heap to a single region on the heap and compresses and frees memory in the process. This evacuation can be performed in parallel on a multi-processor to reduce pause times and increase throughput.
[0122] Example phases within the GC cycle of a concurrent compacting garbage collector include:
[0123] · Phase 1: Identify the objects referred to by the root references (this is not concurrent with the application being executed)
[0124] · Phase 2: Mark the objects reachable from the objects referred to by the root references (this can be concurrent)
[0125] · Phase 3: Identify the objects that have been modified as part of the execution of the program during Phase 2 (this can be concurrent)
[0126] · Phase 4: Relabel the objects identified in Phase 3 (this is not concurrent)
[0127] · Phase 5: Copy the live objects from the source region to the destination region to reclaim the memory space of the source region (this is not concurrent)
[0128] As another example, a load-barrier collector marks and compresses live objects but lazily remaps references to relocated objects. The load-barrier collector relies on "colors" embedded within the references stored on the heap. The colors represent the GC state and track the progress of GC operations on the references. The colors are captured via metadata stored in a bit of the reference.
[0129] At any given moment, all GC threads 506a-b agree on what color is a "good color" or "good GC state". GC threads 506a-b that load a reference from the heap 530 into the call stack first apply a check to determine whether the current color of the reference is good. Similarly, application threads 508a-b that load a reference from the heap 530 into the call stack first apply a check to determine whether the current color of the reference is good. The check can be referred to as a "load barrier". References with good colors will hit a fast path that does not incur additional work. Otherwise, the reference will hit a slow path. The slow path involves certain GC operations that bring the reference from the current GC state to the good GC state. The slot in the heap 530 where the reference resides is updated with an alias of the good color to avoid hitting the slow path subsequently (updating to the good color can also be referred to as "self-healing").
[0130] For example, stale references (references to objects that have been concurrently moved during compaction, meaning the address may point to an out-of-date copy of the object, or another object, or even nothing) are guaranteed not to have a good color. Application threads that attempt to load a reference from the heap first execute a load barrier. Via the load barrier, the reference is identified as stale (not having a good color). Consequently, the reference is updated to point to the new location of the object and associated with a good color. The reference with the updated address and good color is stored in the heap. The reference with the updated address can also be returned to the application thread. However, the reference returned to the application thread does not have to include any color.
[0131] In addition to those described above, additional or alternative types of GC processes may be used. Other types of GC processes may also rely on the "color" of the reference, or metadata related to garbage collection stored within the reference.
[0132] In an embodiment, the color is stored with the heap reference, but not with the dereferenceable reference. The term "heap reference" refers to a reference stored on the heap 530. The term "dereferenceable reference" refers to a reference used by the execution engine to access the value of the object being pointed to by the reference. Obtaining the value of the object being pointed to by the reference is referred to as "dereferencing" the reference. A GC thread 506a-b attempting to dereference a reference stored on the heap 530 first loads the reference from the heap 530 onto the call stack of the GC thread 506a-b. An application thread 508a-b attempting to dereference a reference stored on the heap 530 first loads the reference from the heap 530 onto the call stack of the application thread 508a-b. (For example, the application thread loads the reference into a local variable 401 within a frame 400 of the call stack, as described above with reference Figure 4 as described.) Heap references and / or dereferenceable references are generally referred to herein as "references".
[0133] Reference Figure 6 , Figure 6 illustrates a heap reference and a dereferenceable reference according to an embodiment. Depending on the computing environment, a reference may include any number of bits. For example, in an Intel x86-64 machine, a reference has 64 bits.
[0134] In an embodiment, the dereferenceable reference 600 includes a non-addressable portion 602 and an addressable portion 604. The addressable portion 604 defines the maximum address space that the reference 600 can reach. Depending on the hardware system on which the application executes, the non-addressable portion 602 may be required to conform to a canonical format before the reference 600 is dereferenced. If such a requirement is imposed, the hardware system (such as a processor) generates an error when attempting to dereference a non-conforming dereferenceable reference. Thus, the non-addressable portion 602 of the reference 600 is not available for storing any GC-related metadata, such as GC state. For example, in an Intel x86-64 machine, the addressable portion of the reference has 48 bits, and the non-addressable portion has 16 bits. Based on the constraints imposed by the hardware, the reference can reach at most 2 48 unique addresses. The canonical form requires that the non-addressable portion be a sign extension 610 of the value stored in the addressable portion (that is, the high-order bits 48 to 63 must be a copy of the value stored in bit 47).
[0135] As shown, the addressable portion 604 includes an address 606 and optional additional bits 608. The address 606 refers to the address of the object pointed to by the reference 600. The additional bits 608 may be unused. Alternatively, the additional bits 608 may store metadata, which may or may not be related to garbage collection.
[0136] As described above, the dereferenceable reference 600 includes a reference stored on the call stack. Additionally or alternatively, the dereferenceable reference 600 includes a reference embedded within a compiled method stored in the code cache and / or other memory locations. A compiled method is a method that has been converted from a high-level language (such as bytecode) to a low-level language (such as machine code). An application thread may directly access the compiled method within the code cache or other memory location to execute the compiled method. As an example, the compiled method may be generated by the JIT compiler 109 in Figure 1 . As another example, the compiled method may be generated by another component of the virtual machine.
[0137] In an embodiment, the heap reference 650 includes a transient color bit 652, an address bit 606, and optional additional bits 608. The transient color 652 represents a GC state that tracks the progress of GC operations with respect to the reference 650. The color 652 is “transient” because when the reference is loaded from the heap 530 onto the call stack, the color 652 does not need to be retained with the reference. The additional bits 608 may be unused. Alternatively, the additional bits 608 may store metadata, which may or may not be related to garbage collection. In an embodiment, the transient color 652 is stored in the least significant (rightmost) bits of the heap reference 650. For example, the length of the transient color 652 may be two bytes and is stored in bits 0-15 of the heap reference 650.
[0138] In an embodiment, the transient color 652 includes one or more remapping bits 654. In an embodiment, the remapping bits 654 provide an indication of the current relocation phase of each generation of the GC for that generation. In an embodiment, the GC includes two generations (e.g., a young generation and an old generation), and the remapping bits include a number of bits sufficient to describe the current relocation phase of both the young generation and the old generation. For example, the remapping bits may include 4 bits. In an embodiment, the remapping bits 654 are stored in the most significant portion of the transient color 652. For example, when the transient color 652 is stored in bits 0-15 of the heap reference 650, the remapping bits 654 may constitute bits 12-15 of the heap reference 654.
[0139] The transient color 652 may optionally include additional color bits, including one or more marking bits 656, one or more remembered set bits 658, and one or more other bits 660. In an embodiment, the remap bit 654 may indicate the relocation phase of the GC. In a generational GC, the remap bit 654 may indicate the relocation phase for each generation of the GC. The remap bit will be described in more detail below.
[0140] In an embodiment, the marking bits 656 may indicate the marking parity of the GC. In a generational GC, the marking bits 656 may include representations of the marking parity for different generations of the GC. For example, in a GC that includes a young generation and an old generation, the marking bits 656 may include two bits for indicating the marking parity in the young generation and two bits for indicating the marking parity in the old generation. In another example embodiment, the marking bits 656 may include a first set of bits indicating the marking parity of young generation GC operations and a second set of marking bits indicating the parity of full heap GC operations (which may include only the old generation, or both the old generation and the young generation).
[0141] In an embodiment, the remembered set bits 658 may indicate the remembered set phase of the GC. As a specific example, the remembered set bits may be two bits, where a single bit being set indicates the phase of the remembered set. The remembered set bits indicate potential references from the old generation to the young generation.
[0142] In an embodiment, the other bits 660 may be used to represent other characteristics of the GC state. Alternatively, the other bits 660 may not be used. In some embodiments, the number of the other bits 660 may be determined such that the number of bits in the transient color 652 is an integer multiple of a byte (e.g., the number of bits is divisible by 8). For example, the number of bits in the transient color 652 may be 8 bits or 16 bits. In yet another embodiment, the transient color 652 may represent a completely different set of GC states. The transient color 652 may represent GC states used during additional and / or alternative types of GC processes.
[0143] In an embodiment, a GC cycle may include multiple phases. In some embodiments, a GC system may include separate GC cycles for each generation assigned on the heap. For example, the GC system may include a young generation cycle and an old generation cycle. The young generation GC cycle may include the following phases: mark start, concurrent mark, relocating start, concurrent relocate. In some embodiments, the old generation GC cycle is symmetric to the young generation GC cycle and may include the same phases. In some embodiments, each phase is executed concurrently, meaning that one or more application threads 508a, 508b may continue execution during the phase. In other embodiments, one or more of the phases (e.g., mark start, relocating start) may be non-concurrent. All application threads 508a-b must pause during non-concurrent phases (also referred to as "stop-the world pause" or "STW pause"). In some embodiments, a GC cycle (e.g., young generation GC cycle or old generation GC cycle) begins when the objects on the heap assigned to a particular generation exceed a storage threshold or after a particular period of time has elapsed without a GC cycle.
[0144] The following is a detailed discussion of the phases. In addition to what is discussed below, additional and / or alternative operations may be performed in each phase.
[0145] Mark start: During the mark start phase, the GC updates one or more constants (e.g., "good color") by updating the mark parity and / or remembered set parity for the young generation. During mark start, the GC may capture a snapshot of the remembered set data structure.
[0146] Concurrent mark: The GC threads 506a-b perform an object graph traversal to identify and mark all live objects. The GC threads trace the transitive closure through the heap 530, truncating any traversal leading outside the young generation. If a stale reference is found in the heap 530 during this process, the reference is updated with the current address of the object it references. References in the heap 530 are updated to indicate good color.
[0147] Optionally, per-page liveness information (total number and total size of live objects on each memory page) is recorded. The liveness information may be used to select pages for evacuation.
[0148] Mark end: The GC threads 506a-b mark any enqueued objects and trace the transitive closure of the enqueued objects, and confirm that the marking is complete.
[0149] Relocation start: During relocation start, the GC updates one or more constants (e.g., "good color") by at least updating the remap bits. In an embodiment, the GC threads 506a-b select an empty area as the target space. In another embodiment, additional and / or alternative methods may be used to select a target space for relocated objects.
[0150] Concurrent relocation: Marked source space objects can be relocated to the selected target space (in-place compaction may be used in certain cases). Each object that is moved and contains stale pointers into the young generation that is currently being relocated is added to the remembered set. This helps ensure that pointers are remapped subsequently.
[0151] 4. Load and write barriers
[0152] In one or more embodiments, the GC cycle includes one or more concurrent phases. During a concurrent phase, one or more application threads can execute concurrently with one or more GC threads. When an application thread attempts to load a reference from the heap into the call stack, the application thread can execute a reference load barrier. When an application thread attempts to write a reference to the heap, the application thread can execute a reference write barrier.
[0153] Figure 7 An illustration of a reference load barrier according to an embodiment is shown. As shown, the heap 730 includes addresses 00000008, 00000016, ……, 00000048, 00000049, 00000050. The call stack local variables 732 include registers r1, r2, r3. In the example, the reference includes 32 bits. The color of the heap reference can be indicated by bits 0-15. For example, the color can include 4 remap bits (e.g., bits 12-15) for indicating the relocation phases of the young and old generations, 4 mark bits (e.g., bits 8-11) for indicating the mark parity in the young and old generations, 2 remembered set bits (e.g., bits 6-7) for indicating the remembered set parity in the GC, and six other bits (bits 0-5) that may not be used or may store other metadata.
[0154] Regarding the remapping bits, the bits can use an encoding such that exactly one bit out of four remapping bits is set, where the one set bit indicates the relocation phases for both young generation GC operations and full heap GC operations (which can include only the old generation or both the old generation and the young generation). Specifically, the four remapping bits can be represented as a four-digit binary number. For the remapping bits, the value 0001 can indicate that the full heap relocation is in an even phase and the young generation relocation is in an even phase; the value 0010 can indicate that the full heap relocation is in an even phase and the young generation relocation is in an odd phase; the value 0100 can indicate that the full heap relocation is in an odd phase and the young generation relocation is in an even phase; the value 1000 can indicate that the full heap relocation is in an odd phase and the young generation relocation is in an odd phase. Thus, the four possible values including exactly one set bit represent each of the possible combinations of the relocation phases within the old generation and the young generation.
[0155] The GC can also set a shift value that is one higher than the location of the specific bit that is currently set to good color out of the remapping bits. This ensures that the specific bit is the last bit shifted out of the address. For example, assuming the remapping bits are bits 12 - 15, the shift value can be set to a value between 13 and 16, where the value 13 corresponds to bit 12 as the set bit among the remapping bits, the value 14 corresponds to bit 13 as the set bit among the remapping bits, the value 15 corresponds to bit 14 as the set bit among the remapping bits, and the value 16 corresponds to bit 15 as the set bit among the remapping bits. In an embodiment, the shift value changes at least at the start of each new GC relocation phase and can be set using, for example, compiled method entry barrier patching.
[0156] In an embodiment, the referenced address portion may overlap with the color bits, starting with the set bit of the remapping bits. Thus, depending on the location of the set bit in the remapping bits, the referenced address portion can start anywhere between bits 13 and 16. However, any bits included within the overlap are set to zero. Thus, the method requires that the three least significant bits of each address be zero.
[0157] Sample code can include the following:
[0158]
[0159] Based on the code line Person temp1 = new Person(), the application thread creates a new object in heap 730 and the reference temp1 references this new object. The object (referenced by temp1) is of type Person and includes a name field of type String. The object (referenced by temp1) is stored at address "00000008" within heap 730. The name field of the object (referenced by temp1) is stored at address "00000016" within heap 730. The name field is filled with reference 705. Reference 705 includes color 706 and points to address "0048". Thus, address "00000048" includes the value of the name of the object (referenced by temp1), and the value is "TOM".
[0160] Based on the code line String temp2 = temp1.name, the application thread attempts to load reference 705 in the name field of the object referenced by temp1. The application thread hits reference load barrier 710. Reference load barrier 710 includes an instruction to check whether the color 706 of reference 705 includes a remapping bit that matches the current relocation phase of both the young generation and the old generation. Specifically, the instruction determines whether the correct bit among the remapping bits is set.
[0161] To achieve this, a logical right shift bitwise operation is applied to reference 705. The system can shift the reference to the right n times, where n is equal to the shift value set by the GC. Each bit is shifted n places to the right, and n bits with default values are inserted in the leftmost (e.g., most significant) bit. For example, if the canonical form requires all the most significant bits to be 0, the shift operation can insert n 0s into the leftmost bit. Since color 706 is stored in the least significant (rightmost) bit of reference 705, the right shift operation applied to the reference has the effect of removing color bit 706. Additionally, since the remapping bits are stored in the most significant part of the color, the remapping bits are the last one or more bits removed by the right shift operation. Specifically, the shift value set by the GC corresponds to the location of the exact one bit among the remapping bits that is set to the current "good color".
[0162] Then, the system can determine whether the last bit shifted out of the reference is set (e.g., the correct bit indicating the remapping bit is set). For example, in the x86-64 architecture, the system can determine whether the carry flag and the zero flag are set. In the x86-64 architecture, after a bitwise right shift operation, the carry flag is equal to the last bit shifted out of the reference, and the zero flag is set if all bits in the reference are 0 after the shift operation is completed. Thus, when the correct bit of the remapping bit is set, the carry flag is set; when the reference is a reference to a null value (e.g., address 0), the zero flag is set. If the carry flag is not set and the zero flag is not set, the application thread takes the slow path 714. In other cases (e.g., the carry flag is set or the zero flag is set), the application thread takes the fast path 712. In other system architectures, other techniques can be used to determine whether the last bit shifted out of the reference is set.
[0163] The fast path 712 does not necessarily involve any GC operations, such as remapping the reference and / or marking the object as alive. The color 706 has been removed from the reference 705 by the right shift operation. The resulting "00000048" is saved as the reference 707 in the call stack local variable 732, such as at r3. Then, the application thread can dereference the reference 707. The application thread accesses the address indicated by the reference 707, i.e., the address "00000048" within the heap 730. The application thread fetches the value "TOM" at the address "00000048" within the heap 730.
[0164] When the system determines that the application thread should take the slow path, the application thread can select one of the slow paths in the slow path pool. Specifically, the application thread can reload the reference and select a slow path from the slow path pool based on the color 706. For example, the application thread can remap the address indicated by the reference 705. For example, the application can mark the object pointed to by the reference 705 as alive. Then, the application thread can update the color 706 of the reference 705 to a good color. Additionally, as described above, the application thread can remove the color 706 from the reference 705 for storage in the call stack local variable 732. Specifically, the application thread can apply a logical bitwise right shift operation to the reference 705. The system can shift the reference to the right n times, where n is equal to the shift value set by the GC.
[0165] Figure 8 An illustration of a reference write barrier according to an embodiment is shown. As shown, the heap 830 includes addresses 00000008, 00000016,..., 00000024, 00000032,..., 00000048. The call stack local variable 832 includes registers r1, r2, r3. In the example, the reference includes 32 bits. The color of the heap reference can be indicated by bits 0-15.
[0166] The sample code may include the following:
[0167]
[0168] Based on the code line Person temp2 = new Person(), the application thread creates a new object in the heap 830, and the reference temp2 references this new object. The object (referenced by temp2) is of type Person and includes a name field of type String. The object (referenced by temp2) is stored at the address "00000024" within the heap 830. The name field of the object (referenced by temp2) is stored at the address "00000032" within the heap 830. The name field is filled with the reference 805.
[0169] Based on the code line temp2.name = temp3, the application thread attempts to write the reference 807 from the call stack local variable 832 into the heap 830. Specifically, the application thread attempts to write the reference 807 to the address "00000032", which is the location storing the name field of the object referenced by temp2.
[0170] The application thread hits the reference write barrier 810. The reference write barrier 810 includes an instruction to add the color 806 to the reference 807. Specifically, the application thread determines which color is the good color currently based on the current GC phase. Then, the application thread colors the reference 807 with the good color. Coloring the reference 807 with the good color may include: (a) applying a bitwise left shift operation to the reference to shift the reference left by n times (where n is equal to the shift value set by the GC) and inserting n 0s at the least significant bit of the reference; and (b) applying a logical bitwise OR to the result of the left shift with the good color bitmask, which includes the good color set by the GC in the least significant bits (e.g., bits 0 - 15) and 0s in each other bit. The result of the OR is "00488A40". The application thread writes the result "00488A40" to the address "00000032" in the heap 830.
[0171] 5. Lazy Compression
[0172] Figure 9 Illustrated is a set of example operations for lazy compression according to one or more embodiments. Figure 9 One or more of the operations illustrated in may be modified, rearranged, or omitted together. Thus, Figure 9 the specific sequence of operations illustrated in should not be construed as limiting the scope of one or more embodiments. Some of the operations described below may be performed by the garbage collector, and other operations may be performed by the allocator.
[0173] As described above, relocation is a phase of garbage collection in which the garbage collector relocates live data to create contiguous free memory blocks. For example, in a system that uses bump pointer allocation, relocation may be required.
[0174] In an embodiment, the garbage collector determines a source space corresponding to a relocation set (operation 902). Specifically, to select the relocation set, the garbage collector determines which regions of memory contain enough garbage to warrant being scanned during a subsequent relocation phase. The garbage collector may select the relocation set during the time between marking and actual relocation. The garbage collector may further confirm that the region does not include any objects eligible for promotion to the older generation as a condition for membership in the relocation set for that region. In an embodiment, the garbage collector uses a young generation and an older generation, and the relocation set includes regions in the young generation that contain live objects. The source space corresponds to the set of regions in the young generation that are known at this time to still contain one or more live objects.
[0175] In an embodiment, the garbage collector allocates a target space in anticipation of relocating the objects in the source space (operation 904). The garbage collector may allocate a target space equal to or greater than the combined size of the live objects in the source space. Alternatively or additionally, the garbage collector may be configured to perform in-place relocation without requiring pre-allocation of the target space.
[0176] As described herein, eager relocation (i.e., scanning the relocation set and relocating all live objects in a single sweep) incurs significant overhead associated with relocating objects that will no longer be live in the next collection cycle. In an embodiment, instead of performing eager relocation, the garbage collector fills the lazy free list (LFL) with regions from the source space (operation 906). These regions are regions that, while currently containing live objects, may no longer contain any live objects during the next garbage collection.
[0177] In some embodiments, the garbage collector may sort the regions in the LFL in order of decreasing "liveness". That is, the LFL may give preference to those regions that are least likely to still contain live objects during the next collection cycle. The sorting may be based on information obtained during the marking phase regarding how many live bytes are resident in the objects that survived the most recent garbage collection cycle and will not be promoted to the older generation.
[0178] In an embodiment, prior to the next garbage collection, memory needs to be allocated (e.g., for one or more new objects). In some garbage collection methods, the garbage collector maintains only a single free list of memory regions that are confirmed to be free. If an allocation attempt occurs when the confirmed free list is exhausted, the garbage collector is forced to perform garbage collection to free memory for allocation. In contrast, in an embodiment, in response to determining that the confirmed free list is exhausted (operation 908), the allocator further determines whether any regions are available in the LFL (operation 910).
[0179] If no regions are available in the LFL (i.e., the LFL is exhausted), another garbage collection cycle (operation 918) may be required to free memory. However, if the LFL is not exhausted, the allocator selects a destination region in the LFL for allocation (operation 912). For convenience, the allocator may select the least "live" region in the LFL. For each seemingly live object in the destination region, the allocator executes the load barrier associated with that object (operation 914). If the object is still live, the load barrier relocates the object outside the destination region.
[0180] To determine which load barrier(s) to execute, the allocator may use the live information from the marking phase. For example, the marking phase may generate a bitmap for each heap region, having one bit location for each possible object location (e.g., one bit location for every 8 bytes if working with 8-byte aligned objects). The allocator may traverse or "walk" the bitmap of the destination region. At each location, the allocator checks the corresponding bit to determine whether a seemingly live object exists at that location. If a seemingly live object exists at that location, the allocator executes the load barrier for that object. Thus, fully traversing the bitmap and executing the appropriate load barrier(s) evacuates the region of memory for allocation. The allocator may then allocate the region (operation 916).
[0181] In an embodiment, when the garbage collector performs another garbage collection cycle (operation 918), the garbage collector may perform another marking phase, select a new relocation set, and update the LFL accordingly. Regions that were previously in the LFL but not marked in the new collection cycle are no longer live (i.e., do not contain any live objects) and may be moved to the confirmed free list. In some embodiments, only lazy compaction is performed; the garbage collector may never perform eager relocation. Alternatively, certain conditions may trigger eager relocation. For example, the condition that both the confirmed free list and the LFL are exhausted may trigger eager relocation.
[0182] 6. Example Embodiments
[0183] For clarity purposes, detailed examples are described below. The components and / or operations described below should be understood as a specific example that may not apply to some embodiments. Accordingly, the components and / or operations described below should not be construed as limiting the scope of any of the claims.
[0184] Specifically, Figure 10 An example of lazy compression according to one or more embodiments is illustrated. In this example, the memory includes a plurality of regions, R1 to R6, which are assumed to be regions of the young generation. The garbage collector maintains a free list 1002 of regions that have been identified as free. Specifically, at time T1, regions R2 and R5 are identified as free, and regions R1, R3, R4, and R6 remain live, i.e., each region contains at least one live object.
[0185] Between times T1 and T2, the garbage collector performs a collection cycle in which regions R3 and R6 are freed. The garbage collector selects regions R1 and R4 for the next relocation set and adds these regions to the lazy free list (LFL) 1004 without actually relocating these regions.
[0186] Between garbage collection cycles, one or more allocators allocate memory from the young generation, starting from the regions in the confirmed free list 1002. Specifically, between times T2 and T3, the (one or more) allocators allocate regions R2, R3, R5, and R6. On the next allocation attempt, the free list 1002 is exhausted. Thus, between times T3 and T4, the allocator selects region R1 from the LFL 1004 and performs the (one or more) load barriers associated with any seemingly live objects in region R1. In this example, region R1 includes at least one live object, and the at least one live object is relocated from region R1 to the target space 1006, effectively freeing region R1 for allocation.
[0187] At time T5, the allocator has allocated region R1, and region R1 has been removed from the LFL 1004. Between times T5 and T6, another garbage collection cycle occurs, during which the garbage collector scans the young generation and determines that region R4 no longer includes any live objects (e.g., because the region was not marked during the marking phase). Accordingly, the garbage collector moves region R4 from the LFL 1004 to the free list 1002. In this example, the garbage collector does not select any remaining regions for the next relocation set; thus, the LFL 1004 is now empty.
[0188] At Figure 10In the example illustrated in the figure, the allocation granularity is the entire heap region. Alternatively, the allocation unit can be of finer granularity. For example, in the Z Garbage Collector (ZGC), the heap region tends to be 2 megabytes, while the allocations into the heap region tend to be much less than 2 megabytes. Thus, as used herein, the term "region" does not necessarily refer to the entire heap region; it can instead refer to a smaller granularity region for allocation.
[0189] 7. Computer Networks and Cloud Networks
[0190] In one or more embodiments, a computer network provides connectivity between a set of nodes. The nodes can be local to each other and / or remote from each other. The nodes are connected by a set of links. Examples of links include coaxial cables, unshielded twisted pair, copper cables, optical fibers, and virtual links.
[0191] A subset of the nodes implements the computer network. Examples of such nodes include switches, routers, firewalls, and Network Address Translators (NATs). Another subset of the nodes uses the computer network. Such nodes (also referred to as "hosts") can execute client processes and / or server processes. The client processes request computing services (such as executing a specific application and / or storing a specific amount of data). The server processes respond by, for example, executing the requested service and / or returning the corresponding data.
[0192] The computer network can be a physical network, including physical nodes connected by physical links. A physical node is any digital device. A physical node can be a function-specific hardware device, such as a hardware switch, a hardware router, a hardware firewall, or a hardware NAT. Additionally or alternatively, a physical node can be a general-purpose machine configured to execute various virtual machines and / or applications that perform corresponding functions. A physical link is a physical medium that connects two or more physical nodes. Examples of links include coaxial cables, unshielded twisted pair cables, copper cables, and optical fibers.
[0193] The computer network can be an overlay network. An overlay network is a logical network implemented on top of another network (such as a physical network). Each node in the overlay network corresponds to a corresponding node in the underlying network. Thus, each node in the overlay network is associated with an overlay address (to address the overlay node) and an underlying address (to address the underlying node that implements the overlay node). The overlay nodes can be digital devices and / or software processes (such as virtual machines, application instances, or threads). The links connecting the overlay nodes are implemented as tunnels through the underlying network. The overlay nodes at either end of the tunnel view the underlying multi-hop path between them as a single logical link. Tunneling is performed through encapsulation and decapsulation.
[0194] The client can be local to and / or remote from the computer network. The client can access the computer network via other computer networks, such as a private network or the Internet. The client can use a communication protocol, such as the Hypertext Transfer Protocol (HTTP), to convey requests to the computer network. The requests are conveyed via an interface, such as a client interface (e.g., a web browser), a program interface, or an Application Programming Interface (API).
[0195] In one or more embodiments, the computer network provides connectivity between the client and network resources. Network resources include hardware and / or software configured to execute server processes. Examples of network resources include processors, data storage, virtual machines, containers, and / or software applications. Network resources are shared among multiple clients. The clients independently request computing services from the computer network. Network resources are dynamically assigned to requests and / or clients on an as-needed basis. The network resources assigned to each request and / or client can be scaled up or down based on, for example, (a) the computing services requested by a particular client, (b) the aggregated computing services requested by a particular tenant, and / or (c) the aggregated computing services requested by the computer network. Such a computer network can be referred to as a "cloud network".
[0196] In one or more embodiments, a service provider provides a cloud network to one or more end users. Various service models can be implemented via the cloud network, including but not limited to Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). In SaaS, the service provider provides the end user with the ability to use the service provider's applications executed on network resources. In PaaS, the service provider provides the end user with the ability to deploy custom applications to network resources. Custom applications can be created using programming languages, libraries, services, and tools supported by the service provider. In IaaS, the service provider provides the end user with the ability to provision processing, storage, networking, and other basic computing resources provided by network resources. Any arbitrary application, including an operating system, can be deployed on network resources.
[0197] Computer networks can be implemented in various deployments, including but not limited to private clouds, public clouds, and / or hybrid clouds. In a private cloud, network resources are provisioned for exclusive use by a specific group of one or more entities (the term "entity" as used herein refers to a company, organization, individual, or other entity). The network resources can be located locally at the premises of the specific group of entities and / or remote from the premises. In a public cloud, cloud resources are provisioned for multiple entities (also referred to as "tenants" or "customers") that are independent of each other. The computer network and its network resources can be accessed by client devices corresponding to different tenants. Such a computer network can be referred to as a "multi-tenant computer network". Several tenants can use the same specific network resource at different times and / or at the same time. The network resources can be located locally at the premises of the tenants and / or remote from the premises. In a hybrid cloud, the computer network includes a private cloud and a public cloud. The interface between the private cloud and the public cloud allows data and application portability. Data stored at the private cloud and data stored at the public cloud can be exchanged through the interface. Applications implemented at the private cloud and applications implemented at the public cloud can be dependent on each other. Calls can be made from an application at the private cloud to an application at the public cloud (and vice versa) through the interface.
[0198] In one or more embodiments, the tenants of the multi-tenant computer network are independent of each other. For example, the business or operations of one tenant can be separate from the business or operations of another tenant. Different tenants can have different network requirements for the computer network. Examples of network requirements include processing speed, data storage volume, security requirements, performance requirements, throughput requirements, latency requirements, elasticity requirements, quality of service (QoS) requirements, tenant isolation, and / or consistency. The same computer network may need to implement different network requirements requested by different tenants.
[0199] In a multi-tenant computer network, tenant isolation can be implemented to ensure that the applications and / or data of different tenants do not share with each other. Various tenant isolation schemes can be used. Each tenant can be associated with a tenant identifier (ID). Each network resource of the multi-tenant computer network can be tagged with the tenant ID. A tenant can be allowed to access a specific network resource only when the tenant and the specific network resource are associated with the same tenant ID.
[0200] For example, each application implemented by a computer network can be tagged with a tenant ID, and a tenant is only permitted to access a particular application if the tenant and the particular application are associated with the same tenant ID. Each data structure and / or data set stored by the computer network can be tagged with a tenant ID, and a tenant is only permitted to access a particular data structure and / or data set if the tenant and the particular data structure and / or data set are associated with the same tenant ID. Each database implemented by the computer network can be tagged with a tenant ID, and a tenant is only permitted to access the data of a particular database if the tenant and the particular database are associated with the same tenant ID. Each entry in a database implemented by a multi-tenant computer network can be tagged with a tenant ID, and a tenant is only permitted to access a particular entry if the tenant and the particular entry are associated with the same tenant ID. However, the database can be shared by multiple tenants.
[0201] In one or more embodiments, a subscription list indicates which tenants have authorized access to which network resources. For each network resource, a list of tenant IDs of the tenants authorized to access the network resource can be stored. A tenant is only permitted to access a particular network resource if the tenant's tenant ID is included in the subscription list corresponding to the particular network resource.
[0202] In one or more embodiments, network resources corresponding to different tenants (such as digital devices, virtual machines, application instances, and threads) are isolated from tenant-specific overlay networks maintained by a multi-tenant computer network. As an example, packets from any source device in a tenant overlay network can only be transmitted to other devices within the same tenant overlay network. Encapsulation tunnels can be used to prohibit any transmission from a source device on a tenant overlay network to a device in another tenant overlay network. Specifically, a packet received from a source device can be encapsulated within an outer packet. The outer packet is transmitted from a first encapsulation tunnel endpoint (communicating with the source device in the tenant overlay network) to a second encapsulation tunnel endpoint (communicating with the destination device in the tenant overlay network). The second encapsulation tunnel endpoint de-encapsulates the outer packet to obtain the original packet transmitted by the source device. The original packet is transmitted from the second encapsulation tunnel endpoint to the destination device within the same particular overlay network.
[0203] 8. Hardware Overview
[0204] In one or more embodiments, the techniques described herein are implemented by one or more special-purpose computing devices. The (one or more) special-purpose computing devices can be hard-wired to perform the techniques, and / or can include digital electronic devices (such as one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or network processing units (NPUs)) that are persistently programmed to perform the techniques, or can include one or more general-purpose hardware processors programmed to perform the techniques according to program instructions in firmware, memory, other storage, or a combination thereof. Such special-purpose computing devices can also combine custom hard-wired logic, ASICs, FPGAs, or NPUs with custom programming to accomplish the techniques. The special-purpose computing devices can be desktop computer systems, portable computer systems, handheld devices, networking devices, or any other devices that incorporate hard-wired and / or program logic to implement the techniques.
[0205] For example, Figure 11 is a block diagram of a computer system 1100 on which one or more embodiments of the present invention can be implemented. The computer system 1100 includes a bus 1102 or other communication mechanism for communicating information, and a hardware processor 1104 coupled to the bus 1102 for processing information. The hardware processor 1104 can be, for example, a general-purpose microprocessor.
[0206] The computer system 1100 also includes a main memory 1106 (such as random access memory (RAM) or other dynamic storage device) coupled to the bus 1102 to store information and instructions to be executed by the processor 1104. The main memory 1106 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by the processor 1104. When such instructions are stored in a non-transitory storage medium accessible to the processor 1104, the computer system 1100 is presented as a special-purpose machine customized to perform the operations specified in the instructions.
[0207] The computer system 1100 further includes a read-only memory (ROM) 1108 or other static storage device coupled to the bus 1102 to store static information and instructions for the processor 1104. A storage device 1110, such as a magnetic disk or optical disk, is provided and coupled to the bus 1102 to store information and instructions.
[0208] The computer system 1100 can be coupled via a bus 1102 to a display 1112, such as a cathode ray tube (CRT), to display information to a computer user. An input device 1114, including alphanumeric keys and other keys, is coupled to the bus 1102 to convey information and command selections to the processor 1104. Another type of user input device is a cursor control 1116, such as a mouse, trackball, or cursor direction keys, for conveying direction information and command selections to the processor 1104 and for controlling cursor movement on the display 1112. This input device typically has two degrees of freedom in two axes (a first axis (e.g., x) and a second axis (e.g., y)), which allows the device to specify a location in a plane.
[0209] The computer system 1100 can implement the techniques described herein using customized hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, in combination with the computer system 1100, cause the computer system 1100 to be a special-purpose machine or program it to be a special-purpose machine. In one or more embodiments, the techniques herein are performed by the computer system 1100 in response to one or more sequences of one or more instructions contained in the main memory 1106 being executed by the processor 1104. Such instructions can be read into the main memory 1106 from another storage medium, such as the storage device 1110. Execution of the instruction sequence contained in the main memory 1106 causes the processor 1104 to perform the process steps described herein. Alternatively, the hardwired circuitry can be replaced by, or used in combination with, software instructions.
[0210] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage media can include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as the storage device 1110. Volatile media includes dynamic memory, such as the main memory 1106. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape, or any other magnetic data storage media, compact disc read only memory (CD-ROM), any other optical data storage media, any physical media with hole patterns, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge, content addressable memory (CAM), and ternary content addressable memory (TCAM).
[0211] Storage media is different from transmission media but can be used in combination with transmission media. Transmission media participates in transferring information between storage media. For example, transmission media includes coaxial cables, copper wire, and fiber optics, including the lines of the bus 1102. Transmission media can also take the form of acoustic or light waves, such as those generated during radio frequency (RF) and infrared data communications.
[0212] Various forms of media may be involved in carrying one or more sequences of one or more instructions to the processor 1104 for execution. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer may load the instructions into the dynamic memory of the remote computer and send the instructions using a modem over a telephone line or other communication medium. A modem local to the computer system 1100 may receive the data on the telephone line or other communication medium and convert the data into an infrared signal using an infrared transmitter. An infrared detector may receive the data carried in the infrared signal and appropriate circuitry may place the data on the bus 1102. The bus 1102 carries the data to the main memory 1106, from which the processor 1104 retrieves and executes the instructions. The instructions received by the main memory 1106 may optionally be stored on the storage device 1110 before or after being executed by the processor 1104.
[0213] The computer system 1100 also includes a communication interface 1118 coupled to the bus 1102. The communication interface 1118 provides two-way data communication coupling to a network link 1120 connected to a local network 1122. For example, the communication interface 1118 may be an integrated services digital network (ISDN) card, a cable modem, a satellite modem, or a modem providing a data communication connection to a corresponding type of telephone line. As another example, the communication interface 1118 may be a local area network (LAN) card configured to provide a data communication connection to a compatible LAN. A wireless link may also be implemented. In any such implementation, the communication interface 1118 sends and receives electrical, electromagnetic, or optical signals carrying a digital data stream representing various types of information.
[0214] The network link 1120 generally provides data communication through one or more networks to other data devices. For example, the network link 1120 may provide a connection through the local network 1122 to a host 1124 or to data equipment operated by an Internet service provider (ISP) 1126. The ISP 1126 in turn provides data communication services through the global packet data communication network now commonly referred to as the "Internet" 1128. Both the local network 1122 and the Internet 1128 use electrical, electromagnetic, or optical signals carrying a digital data stream. Signals through various networks and signals on the network link 1120 and through the communication interface 1118, which carry digital data to and from the computer system 1100, are example forms of transmission media.
[0215] The computer system 1100 can send messages and receive data, including program code, via one or more networks, network links 1120, and communication interface 1118. In an Internet example, server 1130 can transfer the requested code of an application via Internet 1128, ISP 1126, local network 1122, and communication interface 1118.
[0216] The received code can be executed by processor 1104 when the code is received, and / or stored in storage device 1110 or other non-volatile memory for later execution.
[0217] 9. Miscellaneous; Extensions
[0218] The embodiments relate to a system having one or more devices, the one or more devices including a hardware processor and configured to perform any of the operations described herein and / or recited in any of the following claims.
[0219] In one or more embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more hardware processors, cause any of the operations described herein and / or recited in any of the claims to be performed.
[0220] Any combination of the features and functions described herein may be used in accordance with one or more embodiments. In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary with implementation. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. The sole and exclusive indicator of the scope of the present invention, and what the applicant intends to be the scope of the present invention, is the literal and equivalent scope of the set of claims sought to be protected in this application, in the specific form in which such claims are sought to be protected, including any subsequent corrections.
Claims
1. One or more non-transitory machine-readable media storing instructions that, when executed by one or more processors, cause performance of operations comprising: Selecting, by a garbage collector, multiple regions of memory to include in a relocation set; Populating, by the garbage collector, a lazy free list (LFL) with the multiple regions selected to be included in the relocation set; After populating the LFL: Determining, by an allocator, that a normal free list managed by the garbage collector has been exhausted; In response to determining that the normal free list has been exhausted: Selecting a region from the LFL; Performing, for each of one or more objects marked as live in the region, one or more load barriers respectively associated with the one or more objects, each respective load barrier being configured to relocate the associated object from the region if the associated object remains live; After performing the one or more load barriers: Allocating the region.
2. The one or more non-transitory machine-readable media of claim 1, wherein performing the one or more load barriers determines that none of the one or more objects remain live, such that the region is allocated without the one or more load barriers relocating any live data from the region.
3. The one or more non-transitory machine-readable media of claim 1, the operations further comprising: Determining, by the allocator, that the normal free list managed by the garbage collector has not been exhausted; In response to determining that the normal free list has not been exhausted: Allocating a free region from the normal free list.
4. The one or more non-transitory machine-readable media of claim 1, the operations further comprising: Determining that both the normal free list and the LFL have been exhausted; In response to determining that both the normal free list and the LFL have been exhausted: Triggering a full garbage collection of the multiple regions.
5. The one or more non-transitory machine-readable media of claim 1, the operations further comprising: Determining, as a condition for inclusion in the relocation set, that no live data in the multiple regions is eligible to be promoted from a current generation to an older generation.
6. The one or more non-transitory machine-readable media of claim 1, the operations further comprising: Allocating, by the garbage collector, a target space large enough to hold all live data in the relocation set.
7. The one or more non-transitory machine-readable media of claim 1, wherein the garbage collector populates the LFL with the multiple regions in an order corresponding to the respective liveness of each region.
8. A system comprising: At least one device including one or more hardware processors, The system being configured to perform the operations of any one of claims 1 to 7.
9. A system comprising components that perform the operations of any one of claims 1 to 7.
10. A method comprising the operations of any one of claims 1 to 7.
Citation Information
Patent Citations
Snapshot at the beginning marking in Z garbage collector
US11734171B2
Colorless roots implementation in z garbage collector
US20220374352A1
Write barrier for remembered set maintenance in generational z garbage collector
US20220374353A1