Method and device for object access
By reading GC thread status information from processor registers and optimizing the object access process of application threads with reference information, the problem of slow access speed in the existing technology is solved, more efficient object access is achieved, and application performance is improved.
Patent Information
- Application Number
- CN202010342983.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-26
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-04-26
AI Technical Summary
In the prior art, when an application thread accesses an object, it is necessary to interact with memory through multiple instructions, resulting in slow access speed and affecting application performance.
By reading the status information of the GC thread from the processor's registers and combining the reference information, the application thread's access process to objects is optimized, including using MABR instructions and speculative access to reduce the number of memory interactions.
It significantly improves the speed of access to objects by application threads, improves application performance, and reduces the problem of missing data caches.
Smart Images

Figure CN113553145B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method and device for object access. Background Art
[0002] In a managed memory scenario, the memory includes a heap, which is an area in the memory where objects are stored. Garbage collection (GC) is a mechanism that uses GC threads to recycle orphan objects in the heap to free up memory space. Orphan objects refer to objects in the heap that have no references to other objects.
[0003] GC threads usually include serial GC threads and concurrent GC threads. Because the application thread (also called mutator thread in JAVA application) needs to be in a waiting state when the serial GC thread is working, which will affect the execution of the application thread, so the concurrent GC thread is usually used to recycle the orphaned objects in the heap.
[0004] In order to prevent the objects that the application thread wants to access from being moved away by the parallel GC thread, a read barrier process is used to prevent the application thread from accessing the object. The current read barrier process requires multiple instructions to interact with the memory to complete the corresponding process when the application thread accesses the object, which affects the speed of the application thread's access to the object, thereby reducing the performance of the application. Summary of the invention
[0005] The embodiment of the present application provides a method for object access, which is used to increase the speed of object access by application threads and improve application performance. The embodiment of the present application also provides a corresponding device.
[0006] A first aspect of the present application provides an object access method, comprising: running an application thread and a garbage collection (GC) thread, wherein the application thread is used to access a second object in memory referenced by a first object, and the GC thread is used to manage objects in memory, wherein the objects in memory include a first object and a second object; if the application thread reads reference information in the first object, then state information of the GC thread is read from a register of a processor, wherein the reference information includes address information of the second object and state information of the reference; and according to the state information and reference information of the GC thread, running the application thread to access the second object.
[0007] The object access method provided in the first aspect can be applied to memory hosting technology, and the programming languages involved in the memory hosting technology include JAVA language and dynamic language (for example, GO language), etc. In memory hosting technology, the memory includes an area for storing objects (object instances), which can be called a heap. In order to release the memory space in time, it is necessary to recycle the isolated objects in the heap, and the isolated objects refer to objects that are not referenced by other objects in the heap. The recycling of isolated objects is called garbage collection (GC), which is a mechanism for recycling isolated objects in the heap through GC threads to release memory space. The GC thread is used to manage objects in the heap, and can mark, organize and delete isolated objects in the heap. The objects in the heap serve applications (applications, APPs), and the application thread is a bridge between applications and objects. When the user operates the application to perform a corresponding function, the application will trigger the application thread corresponding to the function to access the objects in the heap. In JAVA applications, the application thread can be called a mutator thread. The application thread usually accesses the second object through the first object, and the second object is referenced by the first object, and the first object references the second object. The reference relationship between the second object and the first object is recorded when the second object and the first object are generated. When the application thread wants to access the second object, it will first read the reference information in the first object, which can also be called "reference". The reference information will be configured in a field in the first object. The application thread can read the reference information from the corresponding field. The address information in the reference information is used to indicate the location where the second object is stored in the heap. The reference status information refers to whether the reference information has been modified, including whether it has not been modified or has been modified. The status information of the GC thread refers to the information of the current working stage of the GC thread, and the working stage of the GC thread includes the marking stage and the sorting stage. If the GC thread is currently working in the marking stage of marking the object, the status information of the GC thread is marking. If the thread is currently working in the sorting stage of sorting the object, the status information of the GC thread is sorting. Of course, if the GC thread has other working stages, or the marking stage or the sorting stage is further divided, the status information of the GC thread can also be represented as information of other stages or specific divisions, and the status information of the GC thread can be represented in the form of a state value. The registers in the processor can be specially set or implemented by using existing free registers. In the first aspect, the status information of the GC thread is read from the register of the processor, which is much faster than reading the status information of the GC thread from the memory. This can significantly improve the access speed of the running application thread to the object and improve the performance of the application.
[0008] In a possible implementation of the first aspect, the state information of the GC thread is read from the register by calling a mask and branch (MABR) instruction, and the MABR instruction is also used to indicate masking of the jump suggestion output by the branch predictor, and to indicate speculative access to the second object according to the address information.
[0009] In this possible implementation, reading the status information of the GC thread from the register through the MABR instruction belongs to a step in the read barrier process. The read barrier process can be understood as a process that the application thread needs to go through when accessing the second object. The read barrier process in the embodiment of the present application may include reading reference information from the first object through a download instruction, reading the status information of the GC thread from the register of the processor through the MABR instruction, and then performing a state comparison, and accessing the object according to the comparison result. This possible implementation allows the application thread to quickly read the status information of the GC thread from the register through the instruction of the MABR instruction. In addition, in this possible implementation, the branch predictor usually gives a jump suggestion before the comparison result of the status information of the GC thread and the reference information is not available, such as: jumping to the area where the second object is located after the application thread enters the repair function to access the second object, so that the application thread can speculatively go to the area to access the second object according to the jump suggestion. The MABR instruction is used to shield the jump suggestion of the branch predictor, so that the application thread can avoid performing a jump operation according to the jump suggestion output by the branch predictor.
[0010] In a possible implementation of the first aspect, the above-mentioned steps: running the application thread to access the second object according to the status information and reference information of the GC thread, includes: before determining the comparison result according to the status information and reference information of the GC thread, running the application thread to the location indicated by the address information according to the MABR instruction, and speculatively accessing the second object.
[0011] In this possible implementation, the second object is accessed speculatively at the location indicated by the address information according to the MABR instruction. Because the comparison result of this possible implementation will most likely instruct the application thread to access the second object at the location indicated by the address information, this speculative method has a higher probability of success, and can further improve the access speed of the application thread to the object while reducing the probability of speculative errors.
[0012] In a possible implementation manner of the first aspect, the method further includes: if the comparison result indicates that the state information of the GC thread corresponds to the referenced state information, determining that the second object belongs to the location indicated by the address information.
[0013] In this possible implementation, the process of comparing the GC status information and the reference information can be performed using a bitwise AND operation. After the comparison result is obtained by bitwise AND, it is determined whether each bit includes a non-0 bit. If each bit is 0, it means that the status information of the GC thread corresponds to the status information of the reference, which means that the reference information does not need to be modified, and the second object belongs to the location indicated by the address information, indicating that the second object does not need to be migrated. In this case, it means that the speculation of accessing the second object to the location indicated by the address information according to the MABR instruction is successful, and the access to the second object can continue.
[0014] In a possible implementation of the first aspect, the method further includes: if the comparison result indicates that the state information of the GC thread does not correspond to the state information of the reference: running the application thread to migrate the second object from the source location indicated by the address information to the destination location, and modifying the address information in the reference information to the address information of the destination location, and modifying the state information of the reference to correspond to the state information of the GC thread; running the application thread to the destination location to access the second object.
[0015] In this possible implementation, the heap usually contains multiple pages. When the GC thread is in the sorting stage, some pages contain very few orphan objects and do not need to be reclaimed by the GC thread. This type of page can be called an allocated region. Some pages contain more orphan objects and need to be reclaimed by the GC thread. This type of page can be called a source region. Some pages basically do not have any objects (basically blank pages). This type of page can be called a destination region. The source position of the second object is usually in the source region, and the destination position is usually in the destination region. The first object is usually in the allocated region. Of course, the pages contained in these regions are not fixed. The pages originally in the allocated region will be divided into the source region because the number of orphan objects contained increases, and the pages originally in the source region will also be divided into the allocated region because the number of orphan objects decreases. If the comparison result includes a non-zero bit, it means that the state information of the GC thread does not correspond to the state information of the reference, which means that the reference information needs to be modified, and the second object should not belong to the source position indicated by the address information. The application thread then enters the repair function, migrates the second object from the source location to the destination area, and records the storage location of the second object in the destination area, which is the destination location, and modifies the address information in the reference information to the address information of the destination location, and modifies the referenced status information to be modified or tidy. The status information of the GC thread corresponds to the status information of the reference, which may be that the current state of the GC thread is tidying, the status information of the reference is modified or tidying, the current state of the GC thread is marked, and the status information of the reference is unmodified or untidy. After the second object is migrated to the destination location, the application thread accesses the second object at the destination location according to the modified address information. In this way, regardless of whether the application object is located at the source location or the destination location, it can be ensured that the application thread has access to the second object.
[0016] In a possible implementation manner of the first aspect, the method further includes: writing the state information of the GC thread into a register.
[0017] In a possible implementation manner of the first aspect, the method further includes: if the phase of the GC thread changes, updating the status information in the register.
[0018] In this possible implementation, the GC thread may work in different states at different time periods. No matter which state the GC thread is working in, the latest state information will be updated to the register when the stage of the GC thread changes. The register only stores the state information of the GC thread. Every time the stage changes, the previous state information will be overwritten by the latest state information. In this way, it can ensure that the application thread can quickly read the most accurate state information of the GC thread.
[0019] The second aspect of the present application provides an object access device, which has the function of implementing the method of the first aspect or any possible implementation of the first aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, for example: a first processing unit, a reading unit, and a second processing unit.
[0020] A third aspect of the present application provides a processor, which includes a register, wherein the register is used to store status information of a garbage collection GC thread, and the status information of the GC thread is used in the above-mentioned first aspect or any possible implementation method of the first aspect.
[0021] A fourth aspect of the present application provides a computer device, which includes at least one processor, a memory, an input / output (I / O) interface, and computer execution instructions stored in the memory and executable on the processor. When the computer execution instructions are executed by the processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.
[0022] In a fifth aspect, the present application provides a computer-readable storage medium storing one or more computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.
[0023] The sixth aspect of the present application provides a computer program product storing one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the processor executes the method of the above-mentioned first aspect or any possible implementation of the first aspect.
[0024] In a seventh aspect, the present application provides a chip system, which includes a processor, and a device for supporting object access to implement the functions involved in the first aspect or any possible implementation of the first aspect. In a possible design, the chip system may also include a memory, and the memory is used to store the necessary program instructions and data of the device for object access. The chip system may be composed of a chip, or may include a chip and other discrete devices.
[0025] Among them, the technical effects brought about by the second aspect and the seventh aspect or any possible implementation method thereof can refer to the technical effects brought about by the first aspect or different possible implementation methods of the first aspect, and will not be repeated here.
[0026] The embodiment of the present application reads the status information of the GC thread from the register of the processor, which is much faster than reading the status information of the GC thread from the memory. This can significantly improve the access speed of the running application thread to the object and improve the performance of the application. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a schematic diagram of the management process of objects in the heap by the GC thread provided in an embodiment of the present application;
[0028] Figure 2 It is a schematic diagram of an embodiment of a method for object access provided by an embodiment of the present application;
[0029] Figure 3 It is a schematic diagram of another embodiment of the method for object access provided by the application embodiment;
[0030] Figure 4A This is a schematic diagram of writing the status information of the GC thread provided by an embodiment of the present application;
[0031] Figure 4B It is another schematic diagram of writing the status information of the GC thread provided by an embodiment of the present application;
[0032] Figure 5 is an exemplary schematic diagram of a method for object access provided by an embodiment of the present application;
[0033] Figure 6 is an example schematic diagram of a read barrier provided in an embodiment of the present application;
[0034] Figure 7 It is a schematic diagram of an embodiment of a device for object access provided by an embodiment of the present application;
[0035] Figure 8 is a schematic diagram of an embodiment of a computer device provided in an embodiment of the present application;
[0036] Fig. 9 It is a schematic diagram of an embodiment of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of the present application, rather than all embodiments. It is known to those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0038] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0039] The embodiment of the present application provides a method for object access, which is used to increase the access speed of application threads to objects and improve the performance of applications. The embodiment of the present application also provides a corresponding device. The following are detailed descriptions.
[0040] In order to facilitate understanding of the embodiments of the present application, some terms involved in the embodiments of the present application are introduced below.
[0041] Heap: An area in memory used to store objects (object instances). If the memory hosting technology uses the JAVA language, the heap is the largest area in the memory managed by the JAVA virtual machine. This area is a memory area shared by all threads. The heap is created when the JAVA virtual machine is started, and almost all objects allocate memory in the heap.
[0042] Garbage collection (GC): A mechanism that recycles isolated objects in the heap through GC threads to free up memory space.
[0043] The GC thread is used to manage objects in the heap. It can mark, organize, and delete isolated objects in the heap.
[0044] The application thread is a bridge between the application (APP) and the object. When the user operates the application to perform a corresponding function, the application will trigger the application thread corresponding to the function to access the object in the heap. In JAVA applications, the application thread can be called a mutator thread.
[0045] Orphan objects: Objects in the heap that have no references to other objects.
[0046] Surviving objects: objects in the heap that have a reference relationship with other objects. In this reference relationship, a first object pointing to a second object indicates that the first object references the second object, which can also be described as the second object being referenced by the first object.
[0047] Read barrier process: The process that an application thread needs to go through when accessing a referenced object. The read barrier process may include reading reference information from the referenced object through a load instruction, reading the status information of the GC thread from the processor's register through a mask and branch (MABR) instruction, and then performing a status comparison, and accessing the object based on the comparison result.
[0048] In order to understand the management process of objects in the heap by the GC thread, the following Figure 1 This section introduces the process.
[0049] Figure 1 This is a schematic diagram of the management process of objects in the heap by the GC thread in an embodiment of the present application.
[0050] like Figure 1 As shown, the management process of the objects in the heap by the GC thread may include an initial phase, a marking phase, and a finishing phase.
[0051] In the initial stage, Figure 1 The heap shown includes six objects, namely, object A, object B, object C, object D, object E, and object F. Among them, object A and object E are isolated objects, neither of which references other objects, nor is referenced by other objects. Object B references object C, or in other words, object C is referenced by object B, and object D references object F, or in other words, object F is referenced by object D.
[0052] In the marking phase, the GC thread can first determine the root object. The root object is an object that can only reference other objects and cannot be referenced by other objects. Figure 1 The root objects in are object B and object D. Then starting from the root object, recursively mark the objects referenced by the root object as live objects. Figure 1 In the example, the surviving objects marked with “√” are object B, object C, object D, and object F.
[0053] In the finishing phase, the GC thread deletes object A and object E, and releases the memory space occupied by these two objects. Because after deleting object A and object E, there is free memory space between object B, object C, object D, and object F, so objects B, object C, object D, and object F can be moved together to obtain a continuous free memory space. After moving, objects B, object C, object D, and object F are renamed as object B', object C', object D', and object F', and the pointing relationship is modified to object B' pointing to object C', and object D' pointing to object F'. This finishing process can also be called a relocation process.
[0054] Combined with the above Figure 1The stage that the GC thread is in can write the status information of the GC thread into the register of the processor. The status information of the GC thread refers to the information of the current working stage of the GC thread. The working stage of the GC thread includes the marking stage and the sorting stage. If the GC thread is in the marking stage, the status information of the marking stage is written into the register. If the GC thread is in the sorting stage, the status information of the sorting stage is written into the register. The status information of different stages can be represented by different status values, such as: the status information of the GC in the marking stage is represented by 0, and the status information of the GC in the sorting stage is represented by 1. Of course, this is just an example, and the representation method of the status information of the GC in different stages is not limited in the embodiments of the present application.
[0055] When the GC thread manages the objects in the heap, the application thread also needs to access the live objects in the heap. The object access method in the embodiment of the present application is described below in conjunction with the accompanying drawings. The method can be applied to any device or equipment involving object access and GC.
[0056] Figure 2 A schematic diagram of an embodiment of a method for object access in an embodiment of the present application.
[0057] like Figure 2 As shown, an embodiment of the object access method provided by the embodiment of the present application includes:
[0058] 101. Run application threads and garbage collection GC threads.
[0059] The application thread is used to access a second object in the memory that is referenced by the first object, and the GC thread is used to manage objects in the memory, where the objects in the memory include the first object and the second object.
[0060] 102. If the application thread reads the reference information in the first object, the state information of the GC thread is read from the register of the processor.
[0061] The reference information is configured in a corresponding field in the first object, and the application thread can read the reference information from the field.
[0062] The reference information includes the address information of the second object and the reference status information. The address information in the reference information is used to indicate the location where the second object is stored in the heap. The reference status information refers to whether the reference information has been modified, including whether it has not been modified or has been modified.
[0063] The status information of the GC thread refers to the information of the current working stage of the GC thread. The working stages of the GC thread include the marking stage and the grooming stage. When the GC thread is in the marking stage, the status information of the GC thread may be a marking status value used to represent marking. When the GC thread is in the marking stage, the status information of the GC thread may be a grooming status value used to represent grooming.
[0064] Optionally, the state information of the GC thread is read from a register by calling a mask branch predictor MABR instruction.
[0065] 103. According to the status information and reference information of the GC thread, the application thread is run to access the second object.
[0066] The embodiment of the present application reads the status information of the GC thread from the register of the processor, which is much faster than reading the status information of the GC thread from the memory. This can significantly improve the access speed of the running application thread to the object, improve the performance of the application, and avoid the problem of data cache miss (D-cache miss).
[0067] The object access method provided in the embodiment of the present application can also be found in Figure 3 To understand. Figure 3 This is a schematic diagram of another embodiment of the object access method provided in the embodiment of the present application. This embodiment is described in the scenario where the method is applied in a JAVA virtual machine.
[0068] like Figure 3 As shown, another embodiment of the object access method provided by the embodiment of the present application includes:
[0069] 201. The execution engine runs the GC thread to manage objects in the heap.
[0070] The step 201 can be referred to Figure 1 Understand the corresponding process in .
[0071] 202. When the execution engine runs the GC thread at different stages of managing objects, the status information of the current stage of the GC is written into the register of the processor.
[0072] The GC thread will update the latest status information to the register when the phase of the GC thread changes. The register only stores the status information of the current phase of the GC thread. Every time the phase changes, the previous status information will be overwritten by the latest status information.
[0073] The writing process can be found in Figure 4A Understand. Figure 4AAs shown, when the GC thread is in the marking phase, the marking state value is written into the register, and when the GC thread is in the finishing phase, the finishing state value is updated into the register. After the finishing phase, if the GC thread executes in a loop, it enters the marking phase again.
[0074] The writing of GC thread status information is not limited to Figure 4A The method shown can also be referred to Figure 4B To understand.
[0075] When the GC thread enters the pause mark start phase, the GC thread writes the mark start value representing the status information of this phase into the register.
[0076] When the GC thread enters the pause mark stop phase, the GC thread updates the mark stop value representing the status information of this phase into the register.
[0077] Both pause mark start and pause mark stop belong to the marking stage. The processing method for reading mark start value and mark stop value is the same as reading the mark status value.
[0078] When the GC thread enters the stage of preparing to start cleaning (or relocation) (pause relocate start), the GC thread updates the relocate start value representing the status information of this stage into the register.
[0079] Pause relocate start belongs to the finishing phase. The processing method of reading the relocate start value is the same as that of reading the finishing status value.
[0080] If the GC thread executes in a loop, it enters the marking start phase again.
[0081] 203. The execution engine runs the application thread to call the read barrier process including the MABR.
[0082] If the application thread wants to access the second object in the heap, and the second object is referenced by the first object, it needs to go through the read barrier process.
[0083] When the compiler compiles Java code that reads object references, it can insert a Read Barrier instruction sequence in the generated code area (codebuffer) to access objects on the heap.
[0084] 204. The execution engine runs the application thread to read the reference information included in the first object through the load instruction in the read barrier process.
[0085] 205. The execution engine runs the application thread to read the GC status information from the register of the processor.
[0086] The GC status information is status information of the current stage of the GC, for example, a marking status value indicating marking or a grooming status value indicating grooming.
[0087] 206. The execution engine runs the application thread to access the second object according to the state information and reference information of the GC.
[0088] This step 206 may include several optional implementation schemes, which are introduced below respectively.
[0089] Solution 1: Access the second object based on the comparison result of the GC status information and the reference information.
[0090] The process of comparing the GC state information and the reference information can be performed by bitwise AND operation. After the comparison result is obtained by bitwise AND, it is determined whether each bit includes a non-zero bit. If each bit is 0, it means that the second object can be directly accessed according to the reference information. If the comparison result includes a non-zero bit, it means that the application thread needs to enter the repair function first and then access the second object.
[0091] If all bits of the comparison result are 0, it means that the GC state information corresponds to the referenced state information in the reference information, and the application thread can access the second object according to the location indicated by the address information in the reference information. The GC thread state information corresponds to the referenced state information, and the current state of the GC thread is trimming, and the referenced state information is modified or trimmed, or the current state of the GC thread is marked, and the referenced state information is not modified or not trimmed.
[0092] If the comparison result includes a non-zero bit, it means that the GC state information does not correspond to the referenced state information in the reference information, and the application thread needs to enter the modification function to migrate the second object from the source location to the destination location. The GC thread state information does not correspond to the referenced state information because the current state of the GC thread is tidying, and the referenced state information is not modified or not tidying.
[0093] When the comparison result includes non-zero bits, the GC thread is usually in the finishing phase. For this scenario, see Figure 5 See the scene diagram shown for understanding.
[0094] like Figure 5As shown, the first object is foo and the second object is bar. The application thread reads the reference information foo.x in foo through the load instruction, and the foo.x includes the address information of the second object and the state information of the reference is not modified or not sorted.
[0095] Figure 5 The allocated region, source region and destination region are involved. These regions are introduced below.
[0096] The heap usually contains multiple pages, some of which contain very few orphan objects and do not need to be reclaimed by GC threads. This type of page can be called an allocation area. Some pages contain more orphan objects and need to be reclaimed by GC threads. This type of page can be called a source area. Some pages basically have no objects (basically blank pages). This type of page can be called a destination area. The source position of the second object is usually in the source area, and the destination position is usually in the destination area. The first object is usually in the allocation area. Of course, the pages contained in these areas are not fixed. Pages that were originally in the allocation area will be divided into the source area because of the increase in the number of orphan objects they contain, and pages that were originally in the source area will also be divided into the allocation area because of the reduction in orphan objects.
[0097] During the marking phase, the GC thread determines the number of orphan objects and surviving objects contained in each page. After the GC thread enters the sorting phase, it divides the above-mentioned allocation area, source area, and destination area according to the number of orphan objects and surviving objects.
[0098] Should Figure 5 In the scenario shown, if the comparison result includes a non-zero bit, the application thread needs to enter the modification function, migrate bar from the source location to the destination area, and record the storage location of the second object in the destination area, which is the destination location. Then, the address information in foo.x is modified to the address information of the destination location, and the reference status information in foo.x is modified to be modified or arranged, and then bar is accessed to the destination area according to the address information of the destination location.
[0099] Solution 2: Access the second object based on the comparison result of the GC status information and the reference information and the jump suggestion of the branch predictor.
[0100] In this solution 2, the corresponding content of the comparison result can be understood by referring to the explanation of the above solution 1. The difference between this solution 2 and solution 1 is that before the comparison result comes out, the branch predictor's jump suggestion is first adopted to jump to the destination area to speculatively access the second object. After the comparison result comes out, if the comparison result indicates that the GC state information corresponds to the reference state information in the reference information, the application thread needs to go to the source area to access the second object. At this time, the application thread jumps back from the destination area to the source area to access the second object.
[0101] Solution 3: Access the second object based on the comparison result of GC status information and reference information and the MABR instruction.
[0102] The MABR instruction is further used to instruct to mask the jump suggestion output by the branch predictor, and to instruct to speculatively access the second object according to the address information.
[0103] In this solution 3, the corresponding content of the comparison result can be understood by referring to the explanation of the above solution 1. The difference between this solution 3 and solution 1 is that before the comparison result comes out, the jump suggestion of the branch predictor is ignored, and instead the application thread is run according to the MABR instruction to speculatively access the second object at the location indicated by the address information.
[0104] In this case, if the GC thread is in the marking phase, the heap has not yet been divided into the above Figure 5 The allocation area, source area and destination area shown in , the application thread accesses the second object according to the location indicated by the address information according to the MABR instruction.
[0105] If the GC thread is in the sorting stage and the reference information has not been modified or sorted, the location indicated by the address information is the location in the source area. According to the MABR instruction, the second object can be accessed at the corresponding location in the source area first. When the comparison result indicates that the GC status information and the reference status information in the reference information do not correspond, the application thread enters the repair function, migrates the second object to the destination area, and records the storage location of the second object in the destination area, which is the destination location, and then modifies the address information in the reference information to the address information of the destination location, and modifies the reference status information to correspond to the status information of the GC thread. At this time, the second object is accessed at the destination location according to the modified address information.
[0106] If the GC thread is in the sorting stage and the reference information has been modified or sorted, the location indicated by the address information is the destination location in the destination area, indicating that the GC thread has migrated the second object to the destination area and accessed the second object in the destination area according to the MABR instruction. When the comparison result comes out, the GC status information indicates that the status information of the reference in the reference information corresponds to the status information of the reference, and there is no problem of application thread access error.
[0107] It can be seen from the solution 3 that, according to the MABR instruction, there is basically no problem of speculation failure when accessing the second object at the location indicated by the address information, and speculative access to the second object first according to the MABR instruction can also speed up the access speed of the second object, further improving the application performance. In the solution 3, when the GC thread is in the marking phase or the GC thread is not working, the problem of inaccurate branch prediction (branch miss) can also be avoided.
[0108] The read barrier process provided in the embodiment of the present application can also be referred to Figure 6 To understand.
[0109] like Figure 6 As shown, the read barrier process includes:
[0110] 301. Download reference information (reference) through a download instruction.
[0111] 302. Read the status information of the GC thread from the register according to the MARB instruction.
[0112] 303. According to the MARB instruction, the second object is speculatively accessed before the comparison result comes out.
[0113] 304. If the comparison result is 0, it means that the speculation is successful.
[0114] 305. If the comparison result is non-zero, the application thread enters the repair function to perform operations of migrating the second object and modifying the reference information.
[0115] 306. After step 305, access the second object.
[0116] Should Figure 6 The specific process of the scheme shown can be understood by referring to the corresponding contents in the above embodiments, and will not be repeated here.
[0117] The engineering staff conducted multiple tests and comparisons using the solution of the present application and the solution of the prior art. The time for executing object access was shortened by at least 9% compared with the prior art, which effectively improved the application performance.
[0118] The object access method is introduced above. The object access device provided in the embodiment of the present application is introduced below with reference to the accompanying drawings.
[0119] like Figure 7 As shown, an embodiment of the object access device provided by the embodiment of the present application includes:
[0120] The first processing unit 401 is used to run an application thread and a garbage collection GC thread. The application thread is used to access a second object in the memory referenced by the first object. The GC thread is used to manage objects in the memory. The objects in the memory include the first object and the second object.
[0121] The reading unit 402 is used to read the state information of the GC thread from the register of the processor if the application thread run by the first processing unit 401 reads the reference information in the first object, where the reference information includes the address information of the second object and the state information of the reference.
[0122] The second processing unit 403 is used to run the application thread to access the second object according to the state information and reference information of the GC thread read by the reading unit 402.
[0123] The embodiment of the present application reads the status information of the GC thread from the register of the processor, which is much faster than reading the status information of the GC thread from the memory. This can significantly improve the access speed of the running application thread to the object, improve the performance of the application, and avoid the problem of data cache miss (D-cache miss).
[0124] Optionally, the state information of the GC thread is read from the register by calling a mask and branch (MABR) instruction, and the MABR instruction is also used to indicate masking of the jump suggestion output by the branch predictor, and to indicate speculative access to the second object according to the address information.
[0125] Optionally, the second processing unit 403 is configured to, before determining a comparison result according to the state information and reference information of the GC thread, run the application thread to speculatively access the second object at the location indicated by the address information according to the MABR instruction.
[0126] Optionally, the second processing unit 403 is further configured to determine that the second object belongs to the location indicated by the address information if the comparison result indicates that the state information of the GC thread corresponds to the referenced state information.
[0127] Optionally, the second processing unit 403 is also used to, if the comparison result indicates that the state information of the GC thread does not correspond to the state information of the reference: then run the application thread to migrate the second object from the source location indicated by the address information to the destination location, and modify the address information in the reference information to the address information of the destination location, and modify the state information of the reference to correspond to the state information of the GC thread; run the application thread to the destination location to access the second object.
[0128] Optionally, the first processing unit 401 is further configured to write the status information of the GC thread into a register.
[0129] Optionally, the first processing unit 401 is further configured to update the status information in the register if the phase of the GC thread changes.
[0130] The above-mentioned object access device 40 can be understood by referring to the embodiment of the above-mentioned object access method part, and will not be described in detail here.
[0131] An embodiment of the present application further provides a processor, and the processor also includes a register, wherein the register is used to store status information of a garbage collection GC thread, and the status information of the GC thread is used in the above-mentioned object access method embodiment.
[0132] Figure 8 , which is a possible logical structure diagram of a computer device 50 provided in an embodiment of the present application. The computer device 50 includes: a processor 501, a communication interface 502, a memory 503, and a bus 504. The processor 501, the communication interface 502, and the memory 503 are interconnected via the bus 504. In the embodiment of the present application, the processor 501 is used to control and manage the actions of the computer device 50. For example, the processor 501 is used to execute Figure 2 Steps 101 to 103 in , and Figure 3 The communication interface 502 is used to support the computer device 50 to communicate. The memory 503 is used to store program codes and data of the computer device 50.
[0133] Among them, the processor 501 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the contents disclosed in this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 504 can be a peripheral component interconnect standard (Peripheral Component Interconnect, PCI) bus or an extended industry standard architecture (Extended Industry Standard Architecture, EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0134] like Fig. 9As shown, a possible logical structure diagram of a computer device 60 provided in an embodiment of the present application is shown. The computer device 60 includes: a hardware layer 601 and a virtual machine (VM) layer 602, and the VM layer may include one or more VMs. The hardware layer 601 provides hardware resources for the VM to support the operation of the VM, and the VM may be the above-mentioned Figure 3 The JAVA virtual machine in the present application can refer to the above Figure 3 The hardware layer 601 includes hardware resources such as a processor, a communication interface, and a memory. The processor includes a register, which is used to store the status information of the garbage collection GC thread.
[0135] In another embodiment of the present application, a computer-readable storage medium is further provided, wherein the computer-readable storage medium stores computer-executable instructions. When at least one processor of the device executes the computer-executable instructions, the device executes the above Figures 1 to 6 Some embodiments describe methods for accessing objects.
[0136] In another embodiment of the present application, a computer program product is also provided. The computer program product includes computer-executable instructions, which are stored in a computer-readable storage medium; at least one processor of the device can read the computer-executable instructions from the computer-readable storage medium, and at least one processor executes the computer-executable instructions so that the device performs the above Figures 1 to 6 Some embodiments describe methods for accessing objects.
[0137] In another embodiment of the present application, a chip system is further provided, the chip system comprising a processor, a device for supporting object access to implement the above Figures 1 to 6 The object access method described in some embodiments. In a possible design, the chip system may also include a memory, which is used to store program instructions and data necessary for the object access device. The chip system may be composed of a chip or may include a chip and other discrete devices.
[0138] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present application.
[0139] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0140] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0141] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0142] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0143] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.
[0144] The above is only a specific implementation of the embodiment of the present application, but the protection scope of the embodiment of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed in the embodiment of the present application, which should be included in the protection scope of the embodiment of the present application. Therefore, the protection scope of the embodiment of the present application should be based on the protection scope of the claims.
Claims
1. A method for object access, characterized in that: include: Running an application thread and a garbage collection (GC) thread, wherein the application thread is used to access a second object in a memory referenced by a first object, and the GC thread is used to manage objects in the memory, wherein the objects in the memory include the first object and the second object; If the application thread reads the reference information in the first object, then read the state information of the GC thread from the register of the processor, the reference information including the address information of the second object and the reference state information; According to the comparison result of the status information of the GC thread and the reference information, the application thread is run to access the second object; wherein the comparison result is used to instruct the second object to be accessed from the source location indicated by the address information of the second object, or the comparison result is used to instruct the application thread to migrate the second object from the source location to the destination location and access the second object from the destination location.
2. The method according to claim 1, characterized in that The state information of the GC thread is read from the register by calling a mask branch predictor MABR instruction, and the MABR instruction is also used to indicate masking of the jump suggestion output by the branch predictor. Before determining a comparison result based on the state information of the GC thread and the reference information, the application thread is run to speculatively access the second object from the source location indicated by the address information of the second object.
3. The method according to claim 1, characterized in that Before running the application thread to access the second object according to the comparison result between the state information of the GC thread and the reference information, the method further includes: Before determining a comparison result according to the state information of the GC thread and the reference information, the application thread is run to jump to the target location to speculatively access the second object according to a jump suggestion output by a branch predictor.
4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: If the comparison result indicates that the state information of the GC thread corresponds to the state information of the reference, it is determined that the second object belongs to the location indicated by the address information.
5. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: If the comparison result indicates that the state information of the GC thread does not correspond to the referenced state information: then running the application thread to migrate the second object from the source location indicated by the address information to the destination location, and modifying the address information in the reference information to the address information of the destination location, and modifying the referenced state information to correspond to the state information of the GC thread; Run the application thread to the destination location to access the second object.
6. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: The status information of the GC thread is written into the register.
7. The method according to claim 6, characterized in that The method further comprises: If the phase of the GC thread changes, the status information in the register is updated.
8. An object access device, characterized in that: include: A first processing unit is configured to run an application thread and a garbage collection (GC) thread, wherein the application thread is configured to access a second object in a memory referenced by a first object, and the GC thread is configured to manage objects in the memory, wherein the objects in the memory include the first object and the second object; a reading unit, configured to read the state information of the GC thread from a register of a processor if the application thread run by the first processing unit reads the reference information in the first object, the reference information including the address information of the second object and the state information of the reference; The second processing unit is used to run the application thread to access the second object according to the comparison result of the status information of the GC thread read by the reading unit and the reference information; wherein the comparison result is used to instruct the second object to be accessed from the source location indicated by the address information of the second object, or the comparison result is used to instruct the application thread to migrate the second object from the source location to the destination location and access the second object from the destination location.
9. The device according to claim 8, characterized in that The state information of the GC thread is read from the register by calling a mask branch predictor MABR instruction, and the MABR instruction is also used to indicate masking of the jump suggestion output by the branch predictor. Before determining a comparison result based on the state information of the GC thread and the reference information, the application thread is run to speculatively access the second object from the source location indicated by the address information of the second object.
10. The device according to claim 9, characterized in that The second processing unit is configured to, before determining a comparison result according to the state information of the GC thread and the reference information, run the application thread to jump to the target location to speculatively access the second object according to a jump suggestion output by a branch predictor.
11. The device according to any one of claims 8 to 10, characterized in that: The second processing unit is further configured to determine that the second object belongs to the location indicated by the address information if the comparison result indicates that the state information of the GC thread corresponds to the referenced state information.
12. The device according to any one of claims 8 to 10, characterized in that: The second processing unit is also used to, if the comparison result indicates that the state information of the GC thread does not correspond to the state information of the reference: run the application thread to migrate the second object from the source location indicated by the address information to the destination location, and modify the address information in the reference information to the address information of the destination location, and modify the state information of the reference to correspond to the state information of the GC thread; run the application thread to the destination location to access the second object.
13. The device according to any one of claims 8 to 10, characterized in that: The first processing unit is further configured to write the status information of the GC thread into the register.
14. The device according to claim 13, characterized in that The first processing unit is further configured to update the status information in the register if the phase of the GC thread changes.
15. A processor, characterized in that: The processor includes a register, and the register is used to store status information of a garbage collection (GC) thread. The status information of the GC thread is used in the method described in any one of claims 1 to 7.
16. A computing device, characterized in that: comprising a processor and a computer-readable storage medium storing a computer program; The processor is coupled to the computer-readable storage medium, and when the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.
17. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
18. A chip system, characterized in that: The method comprises a processor, wherein the processor is called to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
System and method of computer automatic memory management
CN101046755A
Method, device and system used for accessing RAM (Random Access Memory)
CN107038021A