Method for optimizing Java virtual machine heap memory performance based on HBM

By using NUMA architecture in Java virtual machines, different areas of the Java virtual machine heap memory are stored to HBM and DRAM memory nodes, the traditional DRAM cannot meet the bandwidth and speed of Java program heap memory, and the performance of Java programs is improved.

CN120144232APending Publication Date: 2025-06-13ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510114484.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional DRAM cannot meet the requirements of Java program heap memory bandwidth and speed. Although HBM memory provides higher bandwidth, it has limited capacity and high soft error rate.

Method used

Through the non-unified memory access architecture NUMA, HBM memory and DRAM memory are configured as two independently accessed NUMA nodes, and different areas of the Java virtual machine heap memory (young area, elderly area and metaspace) are stored in different memory nodes, young area is stored in HBM memory nodes, and elderly area and metaspace are stored in DRAM memory nodes.

Benefits of technology

Effectively utilize the advantages of HBM and DRAM, optimize the performance of the Java virtual machine heap memory, and improve the execution speed of Java programs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144232A_ABST
    Figure CN120144232A_ABST
Patent Text Reader

Abstract

According to the method for optimizing the heap memory performance of the Java virtual machine based on the HBM, the advantages of the HBM memory and the advantages of the DRAM memory are effectively exerted, the specific data area is distributed to the HBM memory or the DRAM memory in a targeted mode, the heap memory performance of the Java virtual machine can be optimized, and the performance of a Java program can be improved. The method comprises the following steps: (1) configuring an HBM memory and a DRAM memory into two independently accessible NUMA nodes of a computer by using a non-uniform memory access architecture NUMA, namely an HBM memory node and a DRAM memory node; and (2) storing different areas of a Java virtual machine heap memory into different memory nodes, storing data of a young area into an HBM memory node, and storing data of an old area and a meta-space into a DRAM memory node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for optimizing the heap memory performance of a Java virtual machine based on HBM. Background Art

[0002] The Java virtual machine allocates memory for the objects and data of a Java program from the heap memory. Therefore, improving the heap memory allocation and access efficiency of the Java virtual machine can effectively improve the running speed of the Java program. The heap memory of the Java virtual machine is divided into different areas: the young generation, the old generation, and the metaspace. The young generation is allocated to newly created objects. New objects have a short survival period and are frequently created and destroyed, resulting in a high read / write access frequency in the young generation. The old generation is used to store long-lived objects. The creation and recycling frequency of objects is low, resulting in a low memory access frequency in the old generation. The metaspace is mainly used to store constants, fields, etc., and its access frequency is lower than that of the young generation and the old generation. Since the data access frequency in the young generation is the highest and the read / write is frequent, improving the memory access speed and bandwidth of the young generation can effectively improve the heap memory management and access efficiency of the Java virtual machine and accelerate the execution speed of the Java program.

[0003] With the development of technology, traditional DRAM alone cannot meet the requirements of the Java program heap memory bandwidth and speed. High Bandwidth Memory (HBM) can provide a bandwidth 4-8 times higher than that of traditional DRAM through new technologies such as 3D stacking, while consuming less power and space. However, HBM memory also has its own limitations, such as the capacity is generally limited to a few GB and the soft error rate is relatively high. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies existing in the prior art and propose a method for optimizing the heap memory performance of a Java virtual machine based on HBM, which can effectively utilize the respective advantages of HBM memory and DRAM memory, and specifically allocate specific data areas to HBM memory or DRAM memory, so as to optimize the heap memory performance of the Java virtual machine and improve the performance of the Java program.

[0005] The technical solution adopted by the present invention to solve the above problems is: a method for optimizing the heap memory performance of a Java virtual machine based on HBM, characterized in that it includes the following steps:

[0006] (1) Using the Non-Uniform Memory Access (NUMA) architecture to configure HBM memory and DRAM memory as two independently accessible NUMA nodes of the computer, namely the HBM memory node and the DRAM memory node;

[0007] (2) Store different regions of the Java virtual machine heap memory into different memory nodes. Store the data in the young generation into the HBM memory node, and store the data in the old generation and metaspace into the DRAM memory node.

[0008] The method of the present invention for storing different regions of the Java virtual machine heap memory into different memory nodes is as follows: The key functions for applying memory space for object data in different regions of the heap memory are different. Capture the key function names for allocating memory in the young generation, old generation, and metaspace of the Java virtual machine. These key functions will call the system default memory allocation function; intercept and modify the system default memory allocation function so that it can distinguish which specific region of the heap memory the call comes from, and then bind the applied memory space to the corresponding physical memory node according to the principle that the young generation is bound to the HBM memory node and the old generation and metaspace are bound to the DRAM memory node.

[0009] The present invention captures and stores the key function names for memory allocation in the young generation, old generation, and metaspace of the Java virtual machine heap by means of static code analysis or dynamic recording.

[0010] The method of the present invention for intercepting and modifying the system default memory allocation function is as follows: Write an intercept function with the same name as the system default memory allocation function. This function can call the original memory allocation function, and call relevant functions to obtain the current call stack. Perform symbolic resolution on the call stack address to obtain the specific function name, and compare it with the key function names for memory allocation in different regions of the heap memory to determine the region from which the call comes; once it is determined that the call comes from the young generation of the heap memory, bind the applied memory space to the HBM memory node, otherwise, bind the memory to the DRAM memory node.

[0011] When the present invention performs symbolic resolution on the call stack address to obtain the specific function name, it is necessary to cache the information of the call stack.

[0012] The specific method for the present invention to cache the call stack information is as follows: Use the call stack information as the key to store it in a hash table. Use the backtrace function to capture the current call stack, analyze the symbols in the call stack, and judge the region from which the call comes. Write the stack address and the classification result into the hash table as the hash key and value respectively. For subsequent identical stack addresses, directly read from the cached hash table; each time the call stack address is obtained in the subsequent intercept function, first check whether the cached hash table exists. If it exists, directly return the result; otherwise, execute the complete matching path and update the cache.

[0013] After intercepting the default memory allocation function of the present invention's interception system, the interception function is compiled into a shared library; when starting the Java virtual machine, the shared library is loaded using the LD_PRELOAD environment variable to overwrite the memory allocation function in the system default standard library, enabling the Java virtual machine to call the intercepted and modified memory allocation function, binding the data allocation in the young generation to the HBM memory node, and binding the data allocation in the old generation and metaspace to the DRAM memory node.

[0014] The HBM memory described in the present invention serves as node 0 of the non-uniform memory access architecture NUMA, and the DRAM memory serves as node 1 of the non-uniform memory access architecture NUMA.

[0015] Compared with the prior art, the present invention has the following advantages and effects: According to the characteristics of the HBM memory and in combination with the read / write access characteristics of different regions of the Java virtual machine heap memory, the present invention provides a method for optimizing the performance of the Java virtual machine heap memory based on HBM. In the Java virtual machine heap memory, the object data in the young generation has a short survival cycle and frequent read / write accesses, so the memory allocation in the young generation is bound to the HBM; the object data in the old generation and metaspace has a long life cycle and few read / write accesses, so the allocation in the old generation and metaspace is bound to the DRAM; thus effectively leveraging the respective advantages of the HBM and DRAM to optimize the performance of the Java virtual machine heap memory. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic diagram of the main process of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The present invention will be further described in detail below in conjunction with the drawings and through embodiments.

[0018] The embodiments of the present invention include the following steps:

[0019] (1) Using the non-uniform memory access architecture NUMA, configure the HBM memory and the common DRAM memory as two independently accessible NUMA nodes of the computer, namely the HBM memory node and the DRAM memory node.

[0020] Among them, the HBM memory serves as node 0 of the non-uniform memory access architecture NUMA, that is, node0; the DRAM memory serves as node 1 of the non-uniform memory access architecture NUMA, that is, node1. Subsequently, data storage to node 0 or node 1 can be controlled through the NUMA API programming interface.

[0021] (2) Store different regions of the Java virtual machine heap memory into different memory nodes. The object data in the young generation of the Java virtual machine heap memory has a short life cycle and high read / write access frequency. The characteristics of HBM memory are high bandwidth and low latency, which can effectively improve the access efficiency of frequently read / written data. Therefore, store the data in the young generation into the HBM memory node; the data in the old generation of the heap memory has a long life cycle, low access frequency and large data volume, and store it in the large-capacity DRAM memory node; the data access frequency of the metaspace in the heap memory is lower than that of the old generation, but its data requires high reliability. Therefore, the data in the metaspace is suitable to be stored in the more stable DRAM memory node.

[0022] (1) The method of storing different regions of the Java virtual machine heap memory into different memory nodes is as follows: The key functions for applying for memory space for different regions in the heap memory are different. By using static code analysis or dynamic recording methods, capture the key function names for allocating memory in the young generation, old generation, and metaspace of the Java virtual machine heap. These key functions will call the system default memory allocation function to implement the application for memory space and the allocation of physical memory; intercept and modify the system default memory allocation function so that it can distinguish which specific region in the heap memory the call comes from, and then according to the principle that the young generation is bound to the HBM memory node and the old generation and metaspace are bound to the DRAM memory node, the applied memory space can be bound to the corresponding physical memory node, thus realizing storing different regions of the Java virtual machine heap memory into different memory nodes.

[0023] Taking OpenJDK as an example, the memory allocation in the young generation mainly passes through two key functions, allocate_new_tlab and mem_allocate, while the memory allocation in the old generation mainly passes through

[0024] two key functions, G1Allocator::allocate() and PSOldGen::expand, and the metaspace passes through two key functions, Metaspace::allocate and VirtualSpace::expand. Taking the Linux system as an example, the above different key functions will ultimately call the system default mmap function, as shown in the appendix Figure 1 to implement memory allocation. Therefore, as long as the system default mmap function is intercepted and modified so that it can distinguish which specific region in the heap memory the call comes from, and then bind the young generation allocation to the HBM memory node and bind the old generation and metaspace allocations to the DRAM memory node, the applied memory space can be bound to the corresponding physical memory node, realizing the binding of data in different regions to specific physical memory nodes.

[0025] (2) The method of intercepting and modifying the system default memory allocation function is as follows: Write an intercepting function with the same name as the system default memory allocation function. This function can call the original memory allocation function, call relevant functions to obtain the current call stack, perform symbolic parsing on the call stack address to obtain the specific function name, and compare it with the memory allocation key function names in different regions of the heap memory to determine the region where the call originates. Once it is determined that the call originates from the young generation area of the heap memory, bind the allocated memory space to the HBM memory node; otherwise, bind the memory to the DRAM memory node.

[0026] Taking the Linux system as an example, to intercept and modify the mmap function, the specific steps are to write an intercepting function with the same name as the mmap function. This function can receive all the parameters of the mmap function, and implement the call to the original mmap function, obtain and parse the call stack information, determine the region where the current call originates based on the call stack information, and then use the mbind function to bind the memory space allocated by mmap to the HBM memory node or the DRAM memory node according to the source region.

[0027] First, define a pointer to the original mmap function, assign a value to the pointer, and use this function pointer to call the original mmap function to allocate memory space. The sample code is as follows:

[0028]

[0029]

[0030] If there is no error in calling the original mmap function to allocate space, it is necessary to obtain and parse the call stack information to determine which region of the heap memory the call originates from: If it originates from the young generation area, use the mbind function to bind the just-allocated memory space to the HBM memory node (i.e., node 0); if it originates from the old generation area or the metaspace, use the same mbind to bind the just-allocated memory space to the DRAM memory node (i.e., node 1). The specific sample code is as follows:

[0031]

[0032] }else if(region_type == 2){

[0033] / / Bind the old generation memory to DRAM;

[0034] if(mbind(result, length, MPOL_BIND, &dram_node, sizeof(dram_node)*8, 0)!= 0){

[0035] fprintf(stderr, "Failed to bind memory to DRAM: %s\n", strerror(errno));

[0036] } else {

[0037] printf("[DEBUG]Allocated memory for Old Generation bound to DRAM\n");

[0038] }

[0039] } else if (region_type == 3) {

[0040] / / Memory for Metaspace is bound to DRAM;

[0041] if (mbind(result, length, MPOL_BIND, &dram_node, sizeof(dram_node) * 8, 0) != 0) {

[0042] fprintf(stderr, "Failed to bind memory to DRAM: %s\n", strerror(errno));

[0043] } else {

[0044] printf("[DEBUG]Allocated memory for Metaspace bound to DRAM\n");

[0045] }

[0046] } else {

[0047] printf("[DEBUG]Unrecognized memory request, no specific binding applied.\n");

[0048] }

[0049] return result;

[0050] In the above implementation, the mbind function is used to bind the allocated memory space to the corresponding physical memory. The definitions of the parameters of the mbind function are as follows:

[0051] int mbind(void*start, size_t len, int mode, const unsigned long*nodemask, unsigned long

[0052] maxnode, unsigned flags);

[0053] Among them, the start and len parameters are respectively the starting value and size range of the address space allocated by the mmap function. Among them, nodemask represents the memory node to be bound. Note that the nodemask value of the HBM memory corresponding to node0 is 1, while the nodemask value of the DRAM corresponding to node1 is 2.

[0054] In addition, the determine_region_type function is used to obtain and parse the information of the current call stack, and compare it with the key function names of the memory allocations in different regions of the heap memory recorded previously, so as to determine from which region the current call to mmap comes from. The corresponding sample code is as follows:

[0055] int determine_region_type() {

[0056] void* buffer

[10] ;

[0057] / / Obtain the 10 most recent pieces of information of the call stack;

[0058] int num_frames = backtrace(buffer, 10);

[0059] if (num_frames <= 0) {

[0060] return 0; / / Unable to obtain the call stack;

[0061] }

[0062] / / Parse the call stack

[0063] char** stack_symbols = backtrace_symbols(buffer, num_frames);

[0064] if (stack_symbols == NULL) {

[0065] return 0; / / Unable to parse the call stack;

[0066] }

[0067] / /

[0068] int region_type = 0;

[0069] for (int i = 0; i < num_frames; i++) {

[0070] if (strstr(stack_symbols[i], "allocate_new_tlab") || strstr(stack_symbols[i], "mem_allocate")) {

[0071] region_type = 1; / / Young generation;

[0072] break;

[0073] } else if (strstr(stack_symbols[i], "G1Allocator::allocate") || strstr(stack_symbols[i], "PSOldGen::expand")) {

[0074] region_type = 2; / / Old generation;

[0075] break;

[0076] } else if (strstr(stack_symbols[i], "Metaspace::allocate") || strstr(stack_symbols[i], "VirtualSpace::expand")) {

[0077] region_type = 3; / / Metaspace;

[0078] break;

[0079] }

[0080] }

[0081] free(stack_symbols);

[0082] return region_type;

[0083] In this implementation, the backtrace function is used to obtain the last 10 items of data from the call stack, which are then parsed by backtrace_symbols. The function names after parsing are compared with the key functions for memory allocation in each data area of the heap memory to determine the current call source.

[0084] (3) When performing symbol resolution on the call stack address to obtain the specific function name, to reduce the cost of obtaining and parsing the call stack each time, it is necessary to cache the call stack information. The specific method is as follows: Use the call stack information as the key to store in the hash table. Use the backtrace function to capture the current call stack, analyze the symbols in the call stack, determine the area of the call source, and write the stack address and classification result into the hash table as the hash key and value respectively. For subsequent same stack addresses, directly read from the cached hash table first. Each time the call stack address is obtained in the subsequent interception function, first check whether the cached hash table exists. If it exists, directly return the result; otherwise, execute the complete matching path and update the cache.

[0085] (4) After intercepting the system default memory allocation function, compile the interception function into a shared library. The specific compilation command is as follows:

[0086] gcc -fPIC -shared -o libmmap_hook.so mmap_hook.c -ldl -lnuma -g.

[0087] The compiled shared library is named libmmap_hook.so; when starting the Java virtual machine, use the LD_PRELOAD environment variable to load this shared library. The specific startup command is as follows:

[0088] LD_PRELOAD=. / libmmap_hook.so Java -Xms8G -Xmx16G -XX:+UseG1GC MyApplication.

[0089] In this way, the memory allocation function in the system default standard library can be overwritten; enabling the Java virtual machine to call the intercepted and modified memory allocation function, binding the data allocation in the young generation to the HBM memory node, and binding the data allocation in the old generation and metaspace to the DRAM memory node.

[0090] In addition, it should be noted that for the specific embodiments described in this specification, the shapes, names of the components, etc. can be different. The above content described in this specification is only an example of the structure of the present invention. Any equivalent changes or simple changes made according to the structure, features, and principles conceived in this invention patent are included in the protection scope of this invention patent.

Claims

1. A method for optimizing Java virtual machine heap memory performance based on HBM, characterized in that: The following steps are involved: (i) Use the non-uniform memory access architecture NUMA to configure the HBM memory and DRAM memory as two independently accessible NUMA nodes of the computer, namely the HBM memory node and the DRAM memory node; (ii) Different areas of the Java virtual machine heap memory are stored in different memory nodes, data in the young area is stored in the HBM memory node, and data in the old area and metaspace is stored in the DRAM memory node.

2. The method for optimizing Java virtual machine heap memory performance based on HBM according to claim 1, characterized in that: The method for storing different areas of the Java virtual machine heap memory in different memory nodes is as follows: different areas in the heap memory have different key functions for applying for memory space for object data, and the key function names for allocating memory in the young area, old area and metaspace of the Java virtual machine are captured and stored. These key functions will call the system default memory allocation function; intercept and modify the system default memory allocation function so that it can distinguish which specific area of ​​the heap memory the call comes from, and then bind the applied memory space to the corresponding physical memory node according to the principle of binding the young area to the HBM memory node and the old area and metaspace to the DRAM memory node.

3. The method for optimizing Java virtual machine heap memory performance based on HBM according to claim 2, characterized in that: Static code analysis or dynamic recording is used to capture and store key function names of memory allocation in the young area, old area and metaspace of the Java virtual machine heap.

4. The method for optimizing Java virtual machine heap memory performance based on HBM according to claim 2, characterized in that: The method for intercepting and modifying the system default memory allocation function is as follows: write an interception function with the same name as the system default memory allocation function, which can call the original memory allocation function and call related functions to obtain the current call stack, perform symbolic resolution on the call stack address to obtain the specific function name, and compare it with the key function names of memory allocation in different areas of the heap memory to determine the area where the call comes from; once it is determined that the call comes from the young area of ​​the heap memory, the applied memory space is bound to the HBM memory node, otherwise, the memory is bound to the DRAM memory node.

5. The method for optimizing Java virtual machine heap memory performance based on HBM according to claim 4, characterized in that: When performing symbol resolution on the call stack address to obtain the specific function name, the call stack information needs to be cached.

6. The method for optimizing Java virtual machine heap memory performance based on HBM according to claim 5, characterized in that: The specific method of caching call stack information is as follows: store the call stack information as a key into a hash table, use the backtrace function to capture the current call stack, analyze the symbols in the call stack, determine the area where the call comes from, write the stack address and classification result into the hash table as the hash key and value respectively, and then directly read the hash table used as the cache for the same stack address; each time the call stack address is obtained in the subsequent interception function, first search whether the hash table used as the cache exists, and if so, return the result directly; otherwise, execute the complete matching path and update the cache.

7. The method for optimizing Java virtual machine heap memory performance based on HBM according to claim 2, characterized in that: After intercepting the system's default memory allocation function, compile the intercepted function into a shared library; when starting the Java virtual machine, use the LD_PREDLOAD environment variable to load the shared library, overwriting the memory allocation function in the system's default standard library, so that the Java virtual machine calls the intercepted and modified memory allocation function, binds the data allocation of the young area to the HBM memory node, and binds the data allocation of the old area and metaspace to the DRAM memory node.

8. The method for optimizing Java virtual machine heap memory performance based on HBM according to claim 1, characterized in that: The HBM memory is used as the node 0 of the non-uniform memory access architecture NUMA, and the DRAM memory is used as the node 1 of the non-uniform memory access architecture NUMA.