A memory access conflict optimization method for dynamic instruction conversion based on memory partitioning

By adding a secondary virtual page table and adjusting the memory sub-region size in the ARM core system, a mapping relationship table between the computing core and the memory sub-region is established, the frequent memory access conflicts during the dynamic conversion of x86 instructions to ARM instructions in the ARM core system are solved, and the system performance is improved.

CN119847609BActive Publication Date: 2025-05-23北京麟卓信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510315116.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-05-23
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

In ARM multi-core systems, frequent memory access conflicts lead to system performance degradation during the dynamic conversion of x86 instructions to ARM instructions. Especially in large-scale ARM multi-core systems, existing solutions are difficult to effectively deal with complex memory access situations.

Method used

By adding a secondary virtual page table to the ARM multi-core system, dividing the memory area into N memory sub-regions, and adjusting the size of the memory sub-regions according to the load of the computing core, establishing a mapping relationship table between the computing core and the memory sub-regions, and implementing dynamic instruction conversion to optimize memory access conflicts.

Benefits of technology

It effectively reduces the memory access conflicts during the dynamic conversion of x86 instructions to ARM instructions in ARM core system, and improves the overall performance and execution efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119847609B_ABST
    Figure CN119847609B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for optimizing memory access conflicts in dynamic instruction conversion based on memory partitioning. The method includes establishing a first mapping relationship table by allocating memory sub-areas to each computing core according to load conditions in an ARM many-core system, obtaining a second mapping relationship table between threads and computing cores when loading and executing an executable file in a dynamic instruction conversion mode, obtaining the corresponding thread and memory address of the to-be-converted instruction of the memory access instruction class, obtaining the converted first ARM virtual address and the related first address space according to the first mapping relationship table and the second mapping relationship table, and completing the instruction conversion by determining the adjustment mode of the converted instruction according to the relationship between the first ARM virtual address and the first address space. The method effectively reduces the memory access conflicts in the process of dynamic conversion from x86 instructions to ARM instructions on the ARM many-core system, and improves the execution performance of the program after the dynamic instruction conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of computer architecture and instruction set conversion, and in particular relates to a dynamic instruction conversion memory access conflict optimization method based on memory partitioning. Background Art

[0002] When the ARM many-core system performs the dynamic conversion task from x86 instructions to ARM instructions, many computing cores in the ARM system will access the memory at the same time, which will cause memory access conflicts. Therefore, frequent memory access conflicts will become a key factor affecting system performance. Especially for large-scale ARM many-core systems with more than 64 cores, the existing memory access conflict resolution method is difficult to effectively deal with complex memory access situations, which will cause a significant increase in memory access latency and the overall performance of the ARM system will not meet expectations. Summary of the invention

[0003] In view of this, the present invention provides a memory access conflict optimization method for dynamic instruction conversion based on memory partitioning, which realizes dynamic conversion of x86 instructions to ARM instructions based on dynamic adjustment of computing core partitions.

[0004] The present invention provides a method for optimizing memory access conflicts based on dynamic instruction conversion based on memory partition, which specifically includes the following steps:

[0005] Step 1: Add a secondary virtual page table in the ARM many-core system to realize the mapping of virtual address to virtual address, divide the memory area into N memory sub-areas, where N is the number of computing cores; adjust the memory sub-areas according to the load, bind the computing core and the memory sub-area to establish a first mapping relationship table;

[0006] Step 2: Load and execute the executable file through dynamic instruction conversion, and establish a second mapping relationship table from thread to computing core; if the current instruction to be converted is a memory access instruction, obtain the current thread and the related first memory address, and execute step 3; otherwise, convert it into an ARM instruction and end this process;

[0007] Step 3, obtain the current computing core of the current thread from the second mapping relationship table, obtain the first address space of the current computing core from the first mapping relationship table, and convert the current instruction to be converted into a first ARM instruction and a first ARM virtual address; if the first ARM virtual address is in the first address space, end this process, otherwise execute step 4;

[0008] Step 4: If the first ARM virtual address is not in the secondary virtual page table, modify its base address to the second ARM virtual address, which is the starting address of the first address space, add the first ARM virtual address, the second ARM virtual address and the current computing core to the secondary virtual page table, and end this process; otherwise, execute step 5;

[0009] Step 5. If the computing core of the page table entry corresponding to the first ARM virtual address is the same as the current computing core, then end this process; if they are different and the first ARM instruction is not a memory write operation, then end this process; if they are different and the first ARM instruction is a memory write operation, then convert the first ARM instruction into an instruction sequence consisting of the first ARM instruction and a synchronization barrier instruction and then end this process.

[0010] Furthermore, the method of adjusting the memory sub-area according to the load in step 1 is: obtaining the load of the computing core, determining the allocation weight of the computing core according to the load, and then adjusting its memory sub-area according to the allocation weight.

[0011] Furthermore, the method for obtaining the load of the computing core is: running a load monitoring program on the computing core, regularly obtaining the number of memory accesses, the amount of accessed data and the cache hit rate of the computing core within a set time window, and evaluating the load of the computing core based on these results.

[0012] Furthermore, the memory area is equally divided into N memory sub-areas.

[0013] Furthermore, the method for obtaining the load of the computing core is: evaluating the load of the computing core by counting the number of page fault exceptions and the number of cache misses, wherein the number of page fault exceptions is obtained by modifying the page fault processing function of the ARM kernel, and the number of cache misses is obtained by creating a cache miss event using the kernel performance event interface.

[0014] Furthermore, a central control unit is used to determine the allocation weight of the computing core according to the load, and adjust the memory space of the memory sub-area corresponding to each computing core according to the allocation weight.

[0015] Furthermore, an address translation cache is set in the computing core to save the address mapping relationship between the first ARM virtual address and the second ARM virtual address. When the first ARM virtual address is not in the secondary virtual page table, first, a search is performed in the address translation cache to determine whether there is a relevant record of the first ARM virtual address. If so, the corresponding second ARM virtual address is obtained; otherwise, the base address of the first ARM virtual address is modified to the starting address of the first address space and recorded as the second ARM virtual address, and the address mapping relationship between the first ARM virtual address and the second ARM virtual address is saved in the address translation cache.

[0016] Furthermore, a fast boundary checking algorithm is used to determine the positional relationship between the first ARM virtual address and the first address space.

[0017] Furthermore, the method of establishing the second mapping relationship table from threads to computing cores in step 2 is: starting the monitoring implementation of the thread creation function when the executable file is loaded and executed through dynamic instruction conversion.

[0018] Furthermore, the synchronization barrier instruction in step 5 is a data memory barrier instruction, a data synchronization barrier instruction or an instruction synchronization barrier instruction. Beneficial Effects

[0019] The present invention establishes a first mapping relationship table by allocating memory sub-areas to each computing core according to load conditions in an ARM many-core system, obtains a second mapping relationship table between threads and computing cores when loading and executing an executable file in a dynamic instruction conversion manner, obtains the corresponding thread and memory address for the to-be-converted instruction of the memory access instruction class, obtains the converted first ARM virtual address and the related first address space according to the first mapping relationship table and the second mapping relationship table, and completes the instruction conversion by determining the adjustment method of the converted instruction according to the relationship between the first ARM virtual address and the first address space, thereby effectively reducing the memory access conflict in the dynamic conversion process from x86 instructions to ARM instructions on the ARM many-core system, and improving the execution performance of the program after dynamic instruction conversion. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A flow chart of a method for optimizing memory access conflicts by dynamic instruction conversion based on memory partitioning provided by the present invention. DETAILED DESCRIPTION

[0021] The present invention is described in detail below with reference to the accompanying drawings and with reference to the embodiments.

[0022] The present invention provides a dynamic instruction conversion memory access conflict optimization method based on memory partitioning, the core idea of ​​which is: in an ARM multi-core system, a memory sub-area is allocated to each computing core according to the load condition to establish a first mapping relationship table, a second mapping relationship table between threads and computing cores is obtained when the executable file is loaded and executed in a dynamic instruction conversion manner, for instructions to be converted of the memory access instruction class, the corresponding thread and memory address are obtained, the converted first ARM virtual address and the related first address space are obtained according to the first mapping relationship table and the second mapping relationship table, the adjustment method of the converted instruction is determined according to the relationship between the first ARM virtual address and the first address space, and the instruction conversion is completed.

[0023] The present invention provides a method for optimizing memory access conflicts based on dynamic instruction conversion based on memory partitioning. The specific process is as follows: Figure 1 As shown, the specific steps include:

[0024] Step 1, construct two layers of virtual page tables in the ARM many-core system, including a first-level virtual page table and a second-level virtual page table, wherein the first-level virtual page table is used to realize the mapping of virtual addresses to physical addresses, and the second-level virtual page table is used to realize the mapping of virtual addresses to virtual addresses; the memory area is preliminarily divided into N memory sub-areas according to the number of computing cores in the ARM many-core system, where N is the number of computing cores; the load of each computing core is monitored within a set time window, and the greater the load of the computing core, the higher the allocation weight assigned to it; the memory space of the memory sub-area allocated to each computing core is adjusted according to the allocation weight, and the higher the allocation weight, the larger the allocated memory space; the computing core and the memory sub-area are bound, and a first mapping relationship table containing the binding relationship between the computing core and the memory address space of the memory sub-area is established.

[0025] The initial division of the memory area may be to divide the memory area into N memory sub-areas according to the number of computing cores.

[0026] Furthermore, the present invention implements monitoring of the load of each computing core by running a load monitoring program on each computing core. The load monitoring program regularly counts the number of memory accesses, the amount of accessed data, and the cache hit rate of the computing core within a set time window, and then evaluates the load of the computing core by using these indicators. The load can be evaluated by using indicators such as utilization, process queue length, number of threads and processes, context switching, and number of interrupts.

[0027] In addition, the present invention can also evaluate the load of the computing core by counting the number of page fault exceptions and cache misses. The more these two types of numbers are, the higher the load of memory access is. Specifically, the method for obtaining the number of page fault exceptions is to modify the page fault processing function of the ARM kernel, and the method for obtaining the number of cache misses is to create a cache miss event by using the kernel performance event interface, and the number of cache miss events can be obtained by reading the event information when needed.

[0028] In order to further improve the adjustment efficiency of the computing core load, the present invention implements the adjustment of the load of each computing core based on the central control unit. Specifically, the computing core sends its load to the central control unit through a certain communication method, and the central control unit calculates the allocation weight according to the load of each computing core, and adjusts the memory space of the memory sub-area corresponding to each computing core according to the allocation weight.

[0029] To further improve memory access efficiency, the present invention may also introduce a cache mechanism, that is, an address translation cache ATC is set in each computing core to store the address mapping relationship between the first ARM virtual address and the modified second ARM virtual address.

[0030] Step 2: In the ARM multi-core system, the executable file is loaded and executed through dynamic instruction conversion, the thread creation function is monitored, and the corresponding relationship between the created threads and the computing cores is obtained to establish a second mapping relationship table.

[0031] Step 3, obtain the current instruction to be converted. If the current instruction to be converted is a memory access instruction, obtain the current thread where the current instruction to be converted is located, and record the memory address related to the current instruction to be converted as the first memory address, and then execute step 4; otherwise, convert the current instruction to be converted into an ARM instruction and then execute step 7.

[0032] Step 4. Obtain the computing core where the current thread is located according to the second mapping relationship table and record it as the current computing core. Obtain the memory address space of the memory sub-area bound to the current computing core according to the first mapping relationship table and record it as the first address space. Convert the current instruction to be converted into a first ARM instruction. The memory address in the first ARM instruction is the ARM architecture address of the first memory address after dynamic instruction conversion and is recorded as the first ARM virtual address. If the first ARM virtual address is within the first address space, execute step 7; otherwise, execute step 5.

[0033] In order to improve the efficiency of determining the positional relationship between the first ARM virtual address and the first address space, the present invention adopts a fast boundary checking algorithm.

[0034] Step 5. If the first ARM virtual address does not exist in the secondary virtual page table, the base address of the first ARM virtual address is modified to the starting address of the first address space as the second ARM virtual address, and the first ARM virtual address, the second ARM virtual address and the current computing core are added to the secondary virtual page table, and then step 7 is executed; if the first ARM virtual address exists in the secondary virtual page table, step 6 is executed.

[0035] In order to further improve the modification efficiency of the ARM virtual address, when the first ARM virtual address does not exist in the secondary virtual page table, first check whether there is a relevant record of the first ARM virtual address in the address translation cache of the computing core. If so, obtain the second ARM virtual address corresponding to the first ARM virtual address; otherwise, modify the base address of the first ARM virtual address to the starting address of the first address space as the second ARM virtual address, and save the address mapping relationship between the first ARM virtual address and the second ARM virtual address in the address translation cache.

[0036] Step 6. If the computing core of the corresponding page table entry of the first ARM virtual address in the secondary virtual page table is the same as the current computing core, execute step 7; if the computing core of the corresponding page table entry of the first ARM virtual address in the secondary virtual page table is different from the current computing core and the first ARM instruction is not a write memory operation, execute step 7; if the computing core of the corresponding page table entry of the first ARM virtual address in the secondary virtual page table is different from the current computing core and the first ARM instruction is a write memory operation, convert the first ARM instruction into an instruction sequence consisting of the first ARM instruction and a synchronization barrier instruction, and then execute step 7.

[0037] In the present invention, the synchronization barrier instructions used include data memory barrier instructions, data synchronization barrier instructions, instruction synchronization barrier instructions, etc. In addition, the LDXP instructions and STXP instructions provided by the ARMv8 architecture can also be used for implementation.

[0038] Step 7: If the executable file is executed and the conversion is completed, then the process ends; otherwise, proceed to step 3. Example

[0039] In this embodiment, a dynamic instruction conversion memory access conflict optimization method based on memory partitioning provided by the present invention is adopted to realize efficient execution of x86 architecture executable files on an ARM many-core system, including the following steps:

[0040] S1. Preliminary memory partitioning: Based on the number of computing cores and memory usage requirements of the ARM many-core system, the system memory is initially divided into N partitions. For example, assuming that the system memory is 4GB and there are 128 computing cores, the memory can be divided into 128 equal-sized 32MB partitions as the initial setting.

[0041] S2. Load monitoring and aggregation: Each computing core regularly counts its own memory access times, access data volume, cache hit rate and other indicators to evaluate the load, and then sends the load to the central control unit CCU for aggregation.

[0042] Load monitoring: Run the load monitoring program on each computing core, and count the number of memory accesses, the amount of accessed data, and the cache hit rate of the computing core within 1000 clock cycles every 100 clock cycles. These indicators are used to comprehensively evaluate the load of the computing core. For example, if a computing core has 800 memory accesses in 1000 clock cycles and the cache hit rate is only 30%, it indicates that the memory access load of the computing core is high.

[0043] In addition, the memory access load of the computing core can be evaluated from two perspectives: the number of page fault exceptions and the number of cache miss events. The higher the number of these two times, the higher the memory access load. The specific steps are as follows:

[0044] By modifying the kernel's page fault handling function, the number of page fault exceptions is obtained at the operating system kernel layer. The page fault handling function __handle_mm_fault is modified to custom_mm_fault. The number of page fault exceptions handled by the current core is counted in custom_mm_fault. The sample code is as follows:

[0045] / / Define the memory page fault counter on each CPU

[0046] DEFINE_PER_CPU(unsignedlong,page_fault_count);

[0047] / / Customized page fault handling function

[0048] staticintcustom_mm_fault(structvm_fault*vmf)

[0049] {

[0050] / / Get the number of the current processing core

[0051] intcpu=smp_processor_id();

[0052] / / Increase the page fault counter of the corresponding processing core

[0053] per_cpu(page_fault_count,cpu)++;

[0054] / / Call the original page fault handling function

[0055] return__handle_mm_fault(vmf->vma,vmf->address,vmf->flags);

[0056] }

[0057] / / Save the original page fault processing function pointer

[0058] staticint(*original_mm_fault)(structvm_fault*vmf);

[0059] / / Kernel module initialization function

[0060] staticint__initpage_fault_monitor_init(void)

[0061] {

[0062] / / Save the original page fault handling function

[0063] original_mm_fault=__handle_mm_fault;

[0064] / / Replace with a custom page fault handling function

[0065] __handle_mm_fault=custom_mm_fault;

[0066] }

[0067] Get the number of cache miss events at the operating system kernel layer, create a cache miss event through the kernel performance event interface, and read the event information to get the number of cache miss events when needed. The sample code for creating a kernel performance event is as follows:

[0068] staticint__initcache_miss_monitor_init(void)

[0069] {

[0070] structperf_event_attrpe;

[0071] memset(&pe,0,sizeof(structperf_event_attr));

[0072] / / Configure performance event properties

[0073] pe.type=PERF_TYPE_HARDWARE;

[0074] pe.size=sizeof(structperf_event_attr);

[0075] pe.config=PERF_COUNT_HW_CACHE_MISSES;

[0076] pe.disabled=1;

[0077] pe.exclude_kernel=1;

[0078] pe.exclude_hv=1;

[0079] / / Create performance event

[0080] event=perf_event_create_kernel_counter(&pe,target_cpu,NULL,NULL,NULL);

[0081] if (IS_ERR (event)) {

[0082] printk(KERN_ERR"Failedtocreateperfevent\n");

[0083] returnPTR_ERR(event);

[0084] }

[0085] / / Enable performance events

[0086] perf_event_enable(event);

[0087] return0;

[0088] }

[0089] The sample code for reading cache miss event information is as follows, where count is the number of cache misses:

[0090] staticvoid__exitcache_miss_monitor_exit(void)

[0091] {

[0092] u64count;

[0093] / / Read event count

[0094] count=perf_event_read_value(event);

[0095] }

[0096] For example, the memory access load of a CPU computing core is measured as the sum of the number of page fault exceptions and the number of cache miss events.

[0097] S3, Load information aggregation: Each computing core sends its own load information to the central control unit CCU through shared memory or on-chip network. The central control unit is responsible for the aggregation and processing of load information. Compared with the distributed load adjustment method, it can globally consider the load conditions of each computing core and make more reasonable partition adjustment decisions. The distributed method may lead to uneven adjustment due to the lack of a global perspective.

[0098] An example of assembly code that sends load information to shared memory is as follows:

[0099] MOVR8,#SHARED_MEM_BASE; shared memory base address is stored in R8

[0100] MRCp15,0,R0,c0,c0,5; read the core ID and store it in R0

[0101] ADDR9, R8, R0, LSL#3; Calculate the storage location of the computing core in the shared memory

[0102] STRR5,[R9]; Store the value of the memory access count counter in the shared memory

[0103] S4. Adjust partition size: CCU calculates allocation weights based on the load information of each computing core, and reallocates memory partition sizes according to the weights. Computing cores with high loads receive larger partitions.

[0104] Partition resizing decision: The CCU calculates the allocation weight of each core based on the collected load information. For example, if the memory access load metric of core A is 800 and the memory access load metric of core B is 200, then the allocation weight of core A is 0.8 and the allocation weight of core B is 0.2. Memory partition sizes are reallocated based on the allocation weights. Cores with high loads are allocated to larger memory partitions, and cores with low loads are allocated to smaller memory partitions. Assume that after calculation, the 32MB partition originally allocated to core A needs to be increased to 64MB, while the 32MB partition allocated to core B can be reduced to 16MB.

[0105] The following is an example of assembly code for calculating allocation weights and adjusting partition sizes in the CCU:

[0106] MOVR10,#0;Total load counter initialized to 0

[0107] MOVR11, #128; Store the number of cores in R11

[0108] MOVR12,#0; loop counter initialized to 0

[0109] load_sum_loop:

[0110] LDRR13,[R8,R12,LSL#3]; Read the load information of each core from the shared memory

[0111] ADDR10, R10, R13; Accumulated load information

[0112] ADDR12,R12,#1; loop counter plus 1

[0113] CMPR12, R11; Check whether all cores have been traversed

[0114] BNEload_sum_loop; continue looping if the traversal is not completed

[0115] MOVR12,#0;The loop counter is reinitialized to 0

[0116] weight_calc_loop:

[0117] LDRR13,[R8,R12,LSL#3]; Read the load information of each core from the shared memory

[0118] UDIVR14,R13,R10; calculate the distribution weight

[0119] / / Calculate the new partition size based on the allocation weight and update the memory mapping table

[0120] ADDR12,R12,#1; loop counter plus 1

[0121] CMPR12, R11; Check whether all cores have been traversed

[0122] BNEweight_calc_loop; continue looping if traversal is not completed

[0123] The existing fixed partitioning or simple average distribution method cannot adapt to the dynamic changes of core load. The weight-based distribution method proposed in the present invention can flexibly adjust the partition according to the real-time load of the core, thereby improving the utilization rate of memory resources.

[0124] S5. Partition adjustment implementation: CCU sends the new partition size information to each computing core. Each computing core updates its own memory access configuration based on the received information to ensure that subsequent memory accesses are performed within the newly allocated memory partition. Among them, updating one's own memory access configuration means modifying the information of the memory partition bound to the computing core.

[0125] The following is an example of assembly code for the core to receive and update partition information:

[0126] MOVR15,#CCU_COMM_BASE; The base address for communicating with CCU is stored in R15

[0127] LDRR2,[R15]; Receive the new memory partition start address from CCU

[0128] LDRR3,[R15,#4]; Receive new memory partition size from CCU

[0129] S6. Bind the computing core to the memory partition. The specific steps are as follows:

[0130] S6.1. Assign a unique ID to each computing core, ranging from 0 to N-1. For example, for 128 cores, the core IDs range from 0 to 127.

[0131] S6.2. Bind the computing core to the corresponding memory partition according to its ID, that is, bind the core with core ID i to the i-th memory partition.

[0132] The following is an example of assembly code for setting the memory partition bound to the compute core when it starts:

[0133] MRCp15,0,R0,c0,c0,5; read the core ID and store it in R0

[0134] MOVR1,#0x00000000;The starting address of the memory mapping table is stored in R1

[0135] LDRR2,[R1,R0,LSL#2];Read the memory partition start address corresponding to the core ID from the memory mapping table

[0136] MOVR3,#0x02000000; Each partition size is 32MB, stored in R3

[0137] In a multi-core environment, the one-to-one memory partition binding strategy based on computing core ID is simple, efficient and easy to implement. Compared with the existing random allocation or complex algorithm-based binding methods, the ID-based binding method can quickly determine the correspondence between cores and memory partitions at the hardware level, reducing the overhead in the binding process.

[0138] S7, instruction conversion and memory access control, that is, in the dynamic instruction conversion process, after the x86 instructions are converted to ARM instructions, the ARM instructions involving memory access are efficiently controlled to access the corresponding memory partition. The specific steps are as follows:

[0139] S7.1, instruction pre-analysis: first determine whether it is a memory access instruction, then determine the thread where the current instruction stream is located, and determine the computing core where it is located by the thread. The specific steps are as follows:

[0140] S7.1.1. When executing dynamic instruction conversion, track the thread creation function pthread_create, and establish and record the corresponding relationship between the thread and the CPU computing core according to the thread and CPU computing core binding algorithm in the existing dynamic instruction conversion method.

[0141] S7.1.2, after the instruction conversion module receives the x86 instruction, it pre-analyzes the instruction to determine whether it is a memory access instruction, and if it is a memory access instruction, extracts the memory address information in the instruction. For example, for the x86 instruction MOVEAX, [EBX], extract the memory address in EBX.

[0142] The following is an example of assembly code for instruction pre-analysis in the instruction conversion module:

[0143] LDRR16,[INSTRUCTION_BUFFER]; Read x86 instructions from the instruction buffer

[0144] CMPR16, #0x8B; Determine whether it is a MOV register, [memory address] format instruction

[0145] BNEnot_mem_access; if not, jump to non-memory access processing

[0146] / / Extract memory address information

[0147] LDRR17,[R16+2]; Assume the memory address starts at the third byte of the instruction

[0148] / / Subsequent memory access related processing

[0149] not_mem_access:

[0150] By adding a pre-analysis step before dynamic instruction conversion, memory access instructions can be screened out in advance, and subsequent processing can be carried out in a targeted manner, avoiding unnecessary memory access control operations on non-memory access instructions, and improving the overall conversion efficiency. Existing methods usually perform unified memory access control only after instruction conversion is completed, which increases processing time.

[0151] S7.2, partition matching check: obtain the corresponding CPU computing core according to the thread corresponding to the current instruction stream, compare the extracted memory address with the starting address and size of the memory partition bound to the computing core, and determine whether the address is within the bound memory partition.

[0152] The following is an example of assembly code using the fast bounds checking algorithm:

[0153] MOVR6, R2; R2 is the starting address of the bound memory partition, stored in R6

[0154] ADDR7, R2, R3; R3 is the partition size, calculate the partition end address and store it in R7

[0155] CMPR17, R6; compare access address R17 with the memory partition start address

[0156] BLOaddress_out_of_range; if R17 is less than the starting address, jump to the processing

[0157] CMPR17, R7; compare access address R17 with partition end address

[0158] BHIaddress_out_of_range; if R17 is greater than the end address, jump to the process

[0159] address_out_of_range:

[0160] An optimized fast boundary check algorithm is used for partition matching check. Compared with the existing ordinary comparison algorithm, it can complete the matching judgment of address and partition in a shorter time and improve the efficiency of memory access control.

[0161] S8, address adjustment and cache optimization: If the address is not in the bound memory partition, the corresponding offset calculation is performed to point to the corresponding data in the bound memory partition. At the same time, in order to further improve the memory access efficiency, this embodiment introduces a cache mechanism. A small address translation cache (ATC) is set in each computing core to store the address mapping relationship that has been recently adjusted. When a memory access instruction that requires address adjustment is encountered again, first check in the ATC whether there is a corresponding mapping relationship. If so, the mapping relationship in the cache is directly used for address adjustment to avoid repeated calculations. The specific steps are as follows:

[0162] S8.1. If the address is not in the bound memory partition, perform the corresponding offset calculation to point to the corresponding data in the bound partition. The specific steps are as follows:

[0163] In order to avoid directly modifying the address in the instruction, so as to improve the efficiency of the instruction cache and thus improve the efficiency of the entire converted process, the present invention introduces two layers of virtual page tables in the kernel. The kernel standard is a primary virtual page table, that is, a mapping of virtual addresses to physical addresses. The specific implementation steps of the two-layer virtual page table established by the present invention are as follows:

[0164] S8.1.1. Introduce a virtual address to virtual address mapping table before the kernel standard virtual page table, denoted as vaddr_to_vaddr_table.

[0165] S8.1.2. During instruction conversion, if an address that needs to be mapped is encountered, the mapping target is searched in vaddr_to_vaddr_table. If it does not exist, a mapping page table entry is inserted. The page table entry includes a source address and a target address. Both addresses are recorded in pages and identify the ID of the corresponding CPU computing core. If it exists, it is determined whether the ID of the CPU computing core of the existing page table entry is consistent with the core corresponding to the current instruction. If they are consistent, no operation is performed. When this instruction accesses memory later, it will be converted to the address in the corresponding partition based on the two-layer virtual page table through the kernel's address conversion process. If they are inconsistent, and the current instruction is a memory write operation, a memory synchronization instruction is added after the current instruction.

[0166] Since atomic memory copy operations are required, for example, to avoid data contention in a multi-threaded environment, this embodiment uses LDXP (LoadExclusivePair, atomic loading of two adjacent memory locations) and STXP (StoreExclusivePair, atomic storage of two adjacent memory locations) provided by the ARMv8 architecture, specifically:

[0167] LDXPW0,W1,[R1]; atomically load two adjacent 32-bit data from the memory address pointed to by R1 into the W0 and W1 registers

[0168] STXPW2,W0,W1,[R2]; Atomically store the data in the W0 and W1 registers to the memory address pointed to by R2, and W2 is used to return the operation result

[0169] S8.1.3. By adding a mapping entry, this instruction will be converted to the address in the corresponding partition through the kernel's address translation process based on the two-layer virtual page table when accessing memory.

[0170] In the ARM architecture, this is achieved by modifying kernel functions such as virt_to_phys and pte_offset_map.

[0171] S8.2, when encountering a memory access instruction that requires address adjustment again, first check whether there is a corresponding mapping relationship in ATC. If there is, directly use the mapping relationship in the cache to adjust the address to avoid repeated calculation. The sample code is as follows:

[0172] MOVR18,#ATC_BASE; store the address translation cache base address in R18

[0173] LDRR19,[R18,R17,LSL#2]; Check if there is a mapping relationship for R17 address in ATC

[0174] CMPR19,#0; Check whether the mapping relationship is found

[0175] BEQno_cache_hit; if not found, jump to no cache hit processing

[0176] ADDR17, R19; adjust the address using the mapping relationship in the cache

[0177] / / Perform memory access

[0178] LDRR4,[R17]

[0179] Bend_mem_access

[0180] no_cache_hit:

[0181] / / Perform normal address adjustment calculation

[0182] SUBR17,R17,R6; calculate address offset

[0183] ADDR17, R17, R2; adjust the address to the bound partition

[0184] / / Store the new address mapping in ATC

[0185] STRR17,[R18,R17,LSL#2]

[0186] / / Perform memory access

[0187] LDRR4,[R17]

[0188] end_mem_access:

[0189] The address translation cache (ATC) mechanism is introduced to use the cache to store address mapping relationships, reducing repeated address adjustment calculations and improving memory access efficiency.

[0190] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for optimizing memory access conflicts based on dynamic instruction conversion in memory partitioning, characterized in that: The specific steps include: Step 1: Add a secondary virtual page table in the ARM many-core system to implement virtual address to virtual address mapping, and divide the memory area into N memory sub-areas, where N is the number of computing cores; Adjust the memory sub-region according to the load, bind the computing core and the memory sub-region to establish a first mapping relationship table; Step 2: Load and execute the executable file through dynamic instruction conversion, and establish a second mapping relationship table from thread to computing core; if the current instruction to be converted is a memory access instruction, obtain the current thread and the related first memory address, and execute step 3; Otherwise, convert it into ARM instructions to end this process; Step 3, obtain the current computing core of the current thread from the second mapping relationship table, obtain the first address space of the current computing core from the first mapping relationship table, and convert the current instruction to be converted into a first ARM instruction and a first ARM virtual address; if the first ARM virtual address is in the first address space, end this process, otherwise execute step 4; Step 4: If the first ARM virtual address is not in the secondary virtual page table, modify its base address to the second ARM virtual address, which is the starting address of the first address space, add the first ARM virtual address, the second ARM virtual address and the current computing core to the secondary virtual page table, and end this process; otherwise, execute step 5; Step 5. If the computing core of the page table entry corresponding to the first ARM virtual address is the same as the current computing core, then end this process; if they are different and the first ARM instruction is not a memory write operation, then end this process; if they are different and the first ARM instruction is a memory write operation, then convert the first ARM instruction into an instruction sequence consisting of the first ARM instruction and a synchronization barrier instruction and then end this process.

2. The method for optimizing memory access conflicts in dynamic instruction conversion according to claim 1, characterized in that: The method of adjusting the memory sub-area according to the load in step 1 is: obtaining the load of the computing core, determining the allocation weight of the computing core according to the load, and then adjusting its memory sub-area according to the allocation weight.

3. The method for optimizing memory access conflicts in dynamic instruction conversion according to claim 2, characterized in that: The method of obtaining the load of the computing core is: running a load monitoring program on the computing core, regularly obtaining the number of memory accesses, the amount of accessed data and the cache hit rate of the computing core within a set time window, and evaluating the load of the computing core based on these results.

4. The method for optimizing memory access conflicts in dynamic instruction conversion according to claim 1, characterized in that: Divide the memory area into N memory sub-areas.

5. The method for optimizing memory access conflicts in dynamic instruction conversion according to claim 2, characterized in that: The method for obtaining the load of the computing core is: evaluating the load of the computing core by counting the number of page fault exceptions and the number of cache misses, wherein the number of page fault exceptions is obtained by modifying the page fault processing function of the ARM kernel, and the number of cache misses is obtained by creating a cache miss event using the kernel performance event interface.

6. The method for optimizing memory access conflicts in dynamic instruction conversion according to claim 2, characterized in that: A central control unit is used to determine the allocation weight of the computing core according to the load, and adjust the memory space of the memory sub-area corresponding to each computing core according to the allocation weight.

7. The method for optimizing memory access conflicts in dynamic instruction conversion according to claim 1, characterized in that: An address translation cache is set in the computing core to save the address mapping relationship between the first ARM virtual address and the second ARM virtual address. When the first ARM virtual address is not in the secondary virtual page table, first search in the address translation cache to see whether there is a relevant record of the first ARM virtual address. If so, obtain the corresponding second ARM virtual address; otherwise, modify the base address of the first ARM virtual address to the starting address of the first address space as the second ARM virtual address, and save the address mapping relationship between the first ARM virtual address and the second ARM virtual address in the address translation cache.

8. The method for optimizing memory access conflicts in dynamic instruction conversion according to claim 1, characterized in that: A fast boundary checking algorithm is used to determine the positional relationship between the first ARM virtual address and the first address space.

9. The method for optimizing memory access conflicts in dynamic instruction conversion according to claim 1, characterized in that: The method of establishing the second mapping relationship table from threads to computing cores in step 2 is: starting the monitoring implementation of the thread creation function when the executable file is loaded and executed through dynamic instruction conversion.

10. The method for optimizing memory access conflicts in dynamic instruction conversion according to claim 1, characterized in that: The synchronization barrier instruction in step 5 is a data memory barrier instruction, a data synchronization barrier instruction or an instruction synchronization barrier instruction.

Citation Information

Patent Citations

  • Translation lookaside buffer access method and device, equipment and storage medium

    CN116383102A

  • Prefetch instruction conversion optimization method based on memory access mode virtualization

    CN119440626A