A Dynamic Instruction Conversion Memory Access Optimization Method Based on Data Pre-Alignment

By building memory map tables and data information tables in the ARM system, combining data pre-alignment and instruction merging technology, the problem of low memory access efficiency on the ARM system is solved, and efficient memory access optimization and performance improvement is achieved.

CN119759454BActive Publication Date: 2025-06-20北京麟卓信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510239251.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

When dynamically converting x86 programs to ARM multi-core system for execution, due to differences in memory access mode, memory layout and instruction execution methods, memory access efficiency is low, which affects operating performance and may lead to operation failure.

Method used

Using a dynamic instruction conversion method based on data pre-alignment, by constructing a memory map table in the ARM system, the mapping of x86 segment base address to the base address of the ARM page table is realized, and static or dynamic pre-alignment of instructions and variables is performed based on the data information table, and the relevant instructions are finally merged into batch data transmission instructions to optimize memory access.

Benefits of technology

It realizes efficient memory access for running programs across architectures under dynamic instruction conversion, improves the execution performance of programs on the ARM system, and avoids running failures caused by low memory access efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119759454B_ABST
    Figure CN119759454B_ABST
Patent Text Reader

Abstract

The present invention discloses a memory access optimization method for dynamic instruction conversion based on data pre-alignment. By pre-constructing a memory mapping table for the mapping relationship between the x86 segment base address and the ARM page table base address in the ARM system, a data information table is established when loading and executing an executable file in the dynamic instruction conversion mode, and the initialized and uninitialized global variables and static variables are obtained to establish a first variable table and a second variable table respectively. For statically allocated memory, instruction conversion and static pre-alignment of instructions and variables are completed according to the memory mapping table, the data information table, the first variable table and the second variable table. For dynamically allocated memory, instruction conversion and dynamic pre-alignment of instructions and variables are completed according to the memory mapping table. Finally, based on the address continuity of the memory space, related instructions are merged into batch data transfer instructions to further optimize memory access, realizing memory access optimization during instruction conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer software development, and particularly relates to a dynamic instruction conversion memory access optimization method based on data pre-alignment. Background Art

[0002] When dynamically converting an x86 program to an ARM multi-core system for execution, the memory access efficiency becomes one of the key factors affecting the program execution performance. Due to significant differences in memory access patterns, memory layouts, and instruction execution methods between the x86 system and the ARM system, problems such as poor execution performance or even execution failure caused by low memory access efficiency will occur when the program runs across architectures. Summary of the Invention

[0003] In view of this, the present invention provides a dynamic instruction conversion memory access optimization method based on data pre-alignment, which realizes efficient memory access when running a program across architectures in a dynamic instruction conversion manner.

[0004] A dynamic instruction conversion memory access optimization method based on data pre-alignment provided by the present invention specifically includes the following steps:

[0005] Step 1: Construct a memory mapping table from the x86 segment base address to the ARM page table base address in the ARM system, load and execute the executable file through dynamic instruction conversion, form a first variable table from the initialized variables in global variables and static variables, and form a second variable table from the uninitialized variables; establish a data information table to save the instruction name, variable name, and corresponding ARM starting address, data length, and aligned ARM target address;

[0006] Step 2: Parse the current instruction to be converted to obtain the x86 segment base address and the x86 segment internal offset address corresponding to the instruction and operands. If there is a corresponding entry in the memory mapping table, obtain the ARM page table base address; otherwise, allocate an ARM page table base address for the x86 segment base address and add it to the memory mapping table;

[0007] Step 3: Calculate the ARM starting address of the instruction and variable according to the ARM page table base address and the x86 segment internal offset address. If there is a corresponding ARM target address in the data information table, execute Step 4; if there is a corresponding entry, execute Step 5; otherwise, execute Step 6;

[0008] Step 4: When the variable belongs to the first variable table, perform the alignment operation of the instruction and variable according to the ARM target address, and sequentially copy the initial value of the variable to the aligned buffer; when it belongs to the second variable table, perform the alignment operation of the instruction; convert the current instruction to be converted into an ARM instruction with the same function using the ARM target address as the operand;

[0009] Step 5: Calculate the ARM target addresses of the instructions and variables, save them to the data information table, and convert the currently to-be-converted instruction into an ARM instruction with the same function using the ARM target address as the operand;

[0010] Step 6: Construct a first ARM instruction for checking whether the variable address is aligned, a second ARM instruction with the result of the first ARM instruction as the jump condition, and a third ARM instruction for address alignment, and convert the currently to-be-converted instruction into an ARM instruction sequence composed of ARM instructions with the same function, the first ARM instruction, the second ARM instruction, and the third ARM instruction.

[0011] Further, the memory mapping table is stored in a memory space within the ARM system with a relatively low access latency.

[0012] Further, the entries of the memory mapping table are represented as 64-bit data, where the lower 32 bits store the ARM page table base address and the upper 32 bits store the x86 segment base address.

[0013] Further, it further includes: setting the instructions within a certain range backward from the converted ARM instruction or ARM instruction sequence as the instruction optimization area. If there are multiple memory data transfer instructions with consecutive memory spaces as operands in the instruction optimization area, then merge these memory data transfer instructions into an instruction for batch data transfer to the consecutive memory space; otherwise, keep the instructions within the instruction optimization area unchanged.

[0014] Further, it further includes: adding a data prefetch instruction before the converted ARM instruction or ARM instruction sequence.

[0015] Further, it further includes: when the ARM system catches a data abort exception and the exception address corresponds to the data in the first variable table or the second variable table, then save the ARM start address and data length in the exception information to the data information table, and then end the execution of the program.

[0016] Further, when the executable file has a complete symbol table, the data information table is constructed before the executable file is loaded and executed.

[0017] Further, the method for obtaining the memory space with a relatively low access latency is as follows:

[0018] Step 1.1: Initialize the SysTick timer to generate an interrupt every 1 ms;

[0019] Step 1.2: Perform multiple memory read and write operations on the addresses in the memory space to be measured with a set value as the step size. By recording the values of the SysTick timer before and after the operations, calculate the number of clock cycles consumed by the operations to obtain the access time.

[0020] Step 1.3: The address with the shortest access time is the starting address of the memory space with lower access latency.

[0021] Furthermore, the memory space with lower access latency is a tightly coupled memory. Beneficial effects

[0022] In the present invention, a memory mapping table for the mapping relationship from the x86 segment base address to the ARM page table base address is pre - constructed in the ARM system. When loading and executing an executable file in the dynamic instruction conversion mode, a data information table is established, and the initialized and uninitialized global variables and static variables are obtained to establish a first variable table and a second variable table respectively. For statically allocated memory, instruction conversion and static pre - alignment of instructions and variables are completed based on the memory mapping table, the data information table, the first variable table, and the second variable table. For dynamically allocated memory, instruction conversion and dynamic pre - alignment of instructions and variables are completed based on the memory mapping table. Finally, based on the address continuity of the memory space, related instructions are merged into batch data transfer instructions to further optimize memory access, realizing memory access optimization while performing instruction conversion. Description of the drawings

[0023] Figure 1 It is a flowchart of a memory access optimization method for dynamic instruction conversion based on data pre - alignment provided by the present invention. Detailed implementation manners

[0024] The following are embodiments listed in conjunction with the drawings to describe the present invention in detail.

[0025] A memory access optimization method for dynamic instruction conversion based on data pre - alignment provided by the present invention, its core idea is: in the ARM system, a memory mapping table for the mapping relationship from the x86 segment base address to the ARM page table base address is pre - constructed. When loading and executing an executable file in the dynamic instruction conversion mode, a data information table is established, and the initialized and uninitialized global variables and static variables are obtained to establish a first variable table and a second variable table respectively. For statically allocated memory, instruction conversion and static pre - alignment of instructions and variables are completed based on the memory mapping table, the data information table, the first variable table, and the second variable table. For dynamically allocated memory, instruction conversion and dynamic pre - alignment of instructions and variables are completed based on the memory mapping table. Finally, based on the address continuity of the memory space, related instructions are merged into batch data transfer instructions to further optimize memory access, realizing memory access optimization while performing instruction conversion.

[0026] A memory access optimization method for dynamic instruction conversion based on data pre - alignment provided by the present invention, the specific process is as Figure 1 shown, and specifically includes the following steps:

[0027] Step 1: Pre-build a memory mapping table in the ARM system to implement the mapping from the x86 segment base address to the ARM page table base address, and save the memory mapping table in the memory space with lower access latency within the ARM system.

[0028] Among them, the entries of the memory mapping table can be represented by 64-bit data. The lower 32 bits are used to save the ARM page table base address, and the upper 32 bits are used to save the x86 segment base address, thereby simulating the logical segment memory management method in the x86 architecture in the ARM system.

[0029] Specifically, the method for obtaining the memory space with lower access latency within the ARM system is as follows: It is achieved by executing the built access time test program. The execution process of the access time test program includes:

[0030] Step 1.1: Initialize the SysTick timer to generate an interrupt every 1 ms.

[0031] Step 1.2: Perform multiple memory read and write operations on the addresses in the measured memory space with a set value as the step size. By recording the values of the SysTick timer before and after the operations, calculate the number of clock cycles consumed by the operations to obtain the access time.

[0032] Step 1.3: The address with the shortest access time is the starting address of the memory space with lower access latency within the ARM system.

[0033] Step 2: Load and execute the executable file through dynamic instruction conversion in the ARM system. Obtain the initialized global variables and static variables in the executable file to establish the first variable table, and obtain the uninitialized global variables and static variables in it to establish the second variable table; establish a data information table, which is used to save the instruction name and its ARM starting address, data length, and aligned ARM target address, as well as the variable name and its ARM starting address, data length, and aligned ARM target address. Among them, the ARM starting address and the ARM target address are both ARM virtual addresses.

[0034] Among them, if the executable file has a complete symbol table, then the data information table can be built before the executable file is loaded and executed.

[0035] Step 3: Obtain the current instruction to be converted, parse the x86 segment base address and the x86 segment internal offset address corresponding to the instruction and operands of the current instruction to be converted respectively. Use the obtained x86 segment base address to search in the memory mapping table. If there is an entry containing the x86 segment base address, obtain the ARM page table base address corresponding to the x86 segment base address; otherwise, allocate a new ARM page table base address for the x86 segment base address, and add the mapping relationship between the x86 segment base address and the ARM page table base address to the memory mapping table.

[0036] In the x86 architecture, the address where instructions are stored is related to the operating mode of the system and the memory layout of the program. Usually, instructions are stored within the address range corresponding to the code segment. For example, in the real mode of operation, a simple segmented memory management is adopted. The memory is divided into multiple logical segments, each segment is identified by a segment register, and the code segment is pointed to by the code segment register CS. The CS register stores the high 16 bits of the segment base address, with the default low 4 bits being 0. Shifting the value of CS left by 4 bits (equivalent to multiplying by 16) can obtain the base address of the code segment. The offset address of the instruction is provided by the instruction pointer register IP, and the physical address of the instruction in memory can be determined through the combination of CS and IP.

[0037] The storage address of the operand depends on the type of the operand, the operating mode of the system, and the specific requirements of the program. For example, in the real mode of operation, the data segment is pointed to by the data segment register DS. DS stores the high 16 bits of the segment base address, with the low 4 bits defaulting to 0. By shifting the value of DS left by 4 bits, the base address of the data segment can be obtained. Most ordinary data operands are default stored in the data segment, and the physical address of the operand can be located through DS combined with the in-segment offset address.

[0038] Step 4: Based on the ARM page table base address obtained in Step 3 and the x86 in-segment offset address, calculate the ARM starting addresses of the instructions and variables in the currently to-be-converted instruction; according to the instruction name and its ARM starting address, variable name and its ARM starting address, check whether there is a corresponding entry in the data information table. If there is, it indicates that the variable is a statically allocated variable, and when there is an ARM target address in the entry, execute Step 5, and when there is no ARM target address in the entry, execute Step 6; otherwise, it indicates that the variable is a dynamically allocated variable and execute Step 7.

[0039] Step 5: Read the ARM target address of the instruction and the ARM target address of the variable. When the variable belongs to the first variable table, perform the alignment operation of the instruction and the variable according to the ARM target addresses of the instruction and the variable, and then copy the initial value of the variable to the aligned buffer in the order corresponding to the ARM target address; when the variable belongs to the second variable table, only perform the alignment operation of the instruction according to the ARM target address of the instruction; after converting the currently to-be-converted instruction into an ARM instruction with the same function using the ARM target address as the operand, execute Step 8.

[0040] Step 6: Based on the ARM starting addresses of the instructions and variables in the currently to-be-converted instruction, calculate the alignment addresses of the instructions and variables respectively as the ARM target addresses, and save the instruction name, variable name, their corresponding ARM starting addresses, data lengths, and ARM target addresses into the data information table. After converting the currently to-be-converted instruction into an ARM instruction with the same function using the ARM target address as the operand, execute Step 8.

[0041] In order to further improve the reliability of instruction conversion and execution, the present invention adds an exception handling mechanism. When the ARM system catches a data abort exception and the exception address corresponds to the data in the first variable table or the second variable table, the ARM start address and data length in the exception information are saved into the data information table, and then the execution of the program is ended.

[0042] Step 7: Construct a calculation instruction for checking whether the variable address is aligned, which is denoted as the first ARM instruction; construct a jump instruction conditional on the result of the first ARM instruction, which is denoted as the second ARM instruction; construct an instruction group for address alignment, which is denoted as the third ARM instruction; and convert the current instruction to be converted into an ARM instruction sequence composed of ARM instructions with the same function, the first ARM instruction, the second ARM instruction, and the third ARM instruction.

[0043] Step 8: Set the instructions within a certain range backward from the converted ARM instruction or ARM instruction sequence as the instruction optimization area. If there are multiple memory data transfer instructions with consecutive memory spaces as operands in the instruction optimization area, these memory data transfer instructions are merged into an instruction for batch data transfer to the consecutive memory space; otherwise, the instructions within the instruction optimization area remain unchanged.

[0044] In order to further reduce the memory access latency, the present invention can also add a data prefetch instruction before the converted ARM instruction or ARM instruction sequence, and prefetch in advance the memory data that the ARM instruction or ARM instruction sequence may need to process into the cache.

[0045] Step 9: If the conversion and execution of the executable file are completed, end this process; otherwise, execute Step 3. Embodiment

[0046] In this embodiment, an optimized memory access method for dynamic instruction conversion based on data pre-alignment provided by the present invention is adopted to realize the optimized execution of the x86 architecture executable file on the ARM system, including the following steps:

[0047] S1. Establish and dynamically update a memory mapping table during the instruction conversion process to ensure the effective conversion of memory addresses. The specific steps are as follows:

[0048] S1.1. Construct the basic architecture of the memory mapping table. In the ARM system, in order to realize the memory address conversion from x86 to ARM, a memory mapping table is created.

[0049] S1.2. Save the memory mapping table in the tightly-coupled memory (TCM).

[0050] Since the memory mapping table is accessed frequently, it is necessary to create the memory mapping table in a specific area of the memory to improve the access performance of the memory mapping table. Tightly-Coupled Memory (TCM) Tightly-coupled memory is closely connected to the CPU and has extremely low access latency, capable of completing data read and write operations within one clock cycle. It is usually used to store code and data with extremely high real-time requirements. By measuring the memory access latency, it is judged that the access latency of the memory space is significantly lower, even nearly an order of magnitude lower. In order to accurately measure the latency, the SysTick timer is used for measurement in this embodiment.

[0051] The example code is as follows:

[0052] / / Define test parameters

[0053] #define TEST_LOOPS 10000

[0054] #define TEST_ADDRESS_START 0x00000000

[0055] #define TEST_ADDRESS_END 0x40000000

[0056] #define TEST_STEP 0x1000

[0057] / / Initialize the SysTick timer

[0058] void SysTick_Init(void) {

[0059] SysTick->LOAD = SystemCoreClock / 1000 - 1 / / / / Interrupt once every 1ms

[0060] SysTick->VAL = 0 / /

[0061] SysTick->CTRL = SysTick_CTRL_CLKSOURCE_Msk | SysTick_CTRL_ENABLE_Msk / /

[0062] }

[0063] / / Read the value of the SysTick timer

[0064] uint32_t SysTick_Read(void) {

[0065] return SysTick->VAL / /

[0066] }

[0067] / / Measure memory access time

[0068] uint32_t measure_access_time(uint32_t address) {

[0069] uint32_t start_time, end_time / /

[0070] volatile uint32_t *ptr = (volatile uint32_t *)address / /

[0071] start_time = SysTick_Read(); / /

[0072] for (int i = 0; / / i < TEST_LOOPS; i++) {

[0073] *ptr = i; / / / / Write operation

[0074] volatile uint32_t temp = *ptr; / / / / Read operation

[0075] }

[0076] end_time = SysTick_Read(); / /

[0077] if (end_time > start_time) {

[0078] return (SysTick->LOAD + 1) - (end_time - start_time); / /

[0079] } else {

[0080] return start_time - end_time; / /

[0081] }

[0082] }

[0083] int main(void) {

[0084] SysTick_Init(); / /

[0085] uint32_t min_time = UINT32_MAX; / /

[0086] uint32_t tcm_address = 0; / /

[0087] for (uint32_t address = TEST_ADDRESS_START; address < TEST_ADDRESS_END; address += TEST_STEP) {

[0088] uint32_t time = measure_access_time(address) / /

[0089] if (time < min_time) {

[0090] min_time = time / /

[0091] tcm_address = address / /

[0092] }

[0093] }

[0094] / / Output the possible TCM address

[0095] printf("Possible TCM start address: 0x%08X\n", tcm_address) / /

[0096] while (1) {

[0097] / / Main loop

[0098] }

[0099] }

[0100] The specific steps include:

[0101] S1.2.1, Initialize the SysTick timer: The SysTick_Init function is used to initialize the SysTick timer to generate an interrupt every 1 ms.

[0102] S1.2.2, Measure the memory access time: The measure_access_time function is used to measure the time spent on multiple memory read and write operations to a specified address. By recording the values of the SysTick timer before and after the operation, the number of clock cycles consumed by the operation is calculated.

[0103] S1.2.3, Traverse the address range: In the main function, starting from TEST_ADDRESS_START, traverse to TEST_ADDRESS_END with TEST_STEP as the step size, and call the measure_access_time function to measure each address.

[0104] S1.2.4. Find the fastest access address: Record the minimum access time min_time and the corresponding address tcm_address, which is the starting address of the TCM (denoted as TCM_START_ADDRESS).

[0105] S1.3. Complete it in the kernel. Search in the memory address page table (TLB) of the kernel for an area within the TCM memory range that is unallocated and has a data length meeting the requirements of the memory mapping table. If such an area exists, construct a page table entry for the TLB in the kernel, set the physical address of this page table entry to point to this area, and use the virtual address as the user-mode address of the memory mapping table. If such an area does not exist, call the ordinary memory allocation function in the user mode to allocate the memory mapping table.

[0106] Each entry of the memory mapping table contains the mapping information between the x86 segment base address and the corresponding ARM page table base address. An example of the partial assembly code for creating and initializing the memory mapping table is as follows:

[0107] LDR r0,=0x80000000

[0108] / / Assume the x86 segment base address is stored in r1 and the ARM target physical address is stored in r2

[0109] STR r1,[r0]

[0110] STR r2,[r0,#4]

[0111] S2. Dynamically update the memory mapping table. When a new x86 segment base address is encountered during runtime, the memory mapping table needs to be dynamically updated. The update process can be implemented through a lookup and update function. First, check if there is already a mapping for this x86 segment base address in the memory mapping table. If not, allocate a new ARM page table base address and add it to the memory mapping table. If it exists, consider the possible address relocation situation and update its corresponding ARM page table base address. The following is an example of the assembly code for looking up and updating the memory mapping table:

[0112] / / r3 stores the x86 segment base address to be searched

[0113] LDR r0,=0x80000000

[0114] lookup_loop:

[0115] LDR r4,[r0]

[0116] CMP r4,r3

[0117] BEQ found

[0118] ADD r0,r0,#8

[0119] BNE lookup_loop / / Not found, allocate a new ARM physical address

[0120] / / Assume r5 is used to store the newly allocated ARM physical address

[0121] STR r3, [r0]

[0122] STR r5, [r0, #4]

[0123] B done

[0124] found:

[0125] / / Found, update operations can be performed

[0126] / / Assume the updated ARM physical address is r6

[0127] STR r6, [r0, #4]

[0128] done:

[0129] Existing address translation methods may only adopt simple linear mapping or fixed offset mapping, lacking flexibility. In the present invention, the memory mapping table adopts a dynamic update mechanism, which can flexibly adjust the mapping relationship according to the actual requirements during program operation, support more complex memory management scenarios, and improve the adaptability and scalability of memory address translation. At the same time, the memory mapping table is stored in an independent memory area and adopts a specific data structure, which is convenient for management and update.

[0130] S3. Perform static pre-alignment and dynamic runtime alignment processing on memory data to ensure the alignment of memory access. The specific steps are as follows:

[0131] S3.1. Static pre-alignment.

[0132] The.data section of the ELF file is used to store initialized global variables and static variables, and the.bss section stores uninitialized global variables and static variables.

[0133] At the initial stage of dynamic instruction conversion, for known global data and static data, perform memory alignment processing on the initialized global data and static data based on the data information table, and perform address offset processing on the corresponding instructions. For uninitialized global data and static data, only perform address offset processing on the corresponding instructions. The specific steps are as follows:

[0134] S3.1.1. When the program is loaded, if the program has a complete symbol table, construct a data information table based on the symbol table. If the data information table is not empty, execute S3.1.2; otherwise, execute S3.1.3.

[0135] S3.1.2. Align the data in the.data or.bss segments based on the data information table. The specific steps are as follows:

[0136] Align the initialized global and static data, and record the target address in the corresponding item. Then, copy the initialized data recorded in the data information table to the aligned buffer in the order of alignment. For the uninitialized global and static data, only record the target address in the corresponding item.

[0137] S3.1.3. If the operand of the current instruction to be converted is data in the.data or.bss segment, find the corresponding aligned address in the data information table according to the original offset of this data; if found, change the operand to this aligned address. If not found, an alignment exception may occur during execution.

[0138] If a data abort exception (DataAbortException) is caught and the exception address is data in the.data or.bss segment, record the start address and data length in the exception information to the data information table, and end the program execution.

[0139] S3.2. Dynamic runtime alignment.

[0140] For the memory dynamically allocated during runtime, perform dynamic alignment checks and adjustments during each memory access. When using the LDR or STR instructions, check whether the address is aligned through the AND instruction. If not aligned, adjust the address. In addition, for unaligned memory accesses, padding or reorganization operations can be adopted to ensure aligned access. The following is an example of the assembly code for dynamic runtime alignment:

[0141] / / Assume the data address is stored in r4

[0142] LDR r5,[r4]

[0143] / / Check if it is 4-byte aligned

[0144] AND r6,r4,#0x3

[0145] CMP r6,#0

[0146] BNE unaligned

[0147] / / The address is aligned, continue processing

[0148] B done

[0149] unaligned:

[0150] / / Not aligned, make adjustments

[0151] ADDr4,r4,#4

[0152] LDRr5,[r4]

[0153] / / Process unaligned data, may need to reorganize according to data type

[0154] / / Assume the data is a 32-bit integer, can combine the front and back bytes

[0155] LDRBr7,[r4,#-1]

[0156] LDRBr8,[r4,#-2]

[0157] LDRBr9,[r4,#-3]

[0158] LSLr7,r7,#24

[0159] LSLr8,r8,#16

[0160] LSLr9,r9,#8

[0161] ORRr5,r5,r7

[0162] ORRr5,r5,r8

[0163] ORRr5,r5,r9

[0164] done:

[0165] Different from the existing method which only performs memory alignment processing at compile time, the dynamic runtime alignment method of the present invention can handle the alignment problem of dynamically allocated memory during runtime. At the same time, it combines data reorganization technology, can effectively handle the access of unaligned data, avoids the performance overhead caused by unalignment, and ensures the efficiency of memory access and data integrity.

[0166] S4. On the basis of pre-alignment, optimize memory access instructions, select the optimal instructions and perform prefetch optimization.

[0167] When converting x86 memory access instructions to ARM instructions, for consecutive memory accesses, traverse the starting address and data length of the memory accessed by consecutive LDR / STR instructions. If the concatenated address space is consecutive, merge the memory regions accessed by consecutive LDR / STR instructions, and use LDM and STM instructions to batch load and store data instead of using LDR and STR instructions multiple times. The following is an example of assembly code for batch data processing using LDM and STM:

[0168] / / Assume the data address range to be loaded starts from r0 and the data length is 16 bytes

[0169] Load Multiple Register r0, {r1 - r4}

[0170] / / Store data

[0171] Store Multiple Register r5, {r6 - r9}

[0172] For different data types and sizes, select appropriate load and store instructions, such as LDRB for bytes, LDRH for half - words, and LDR for words.

[0173] S5, Prefetch optimization.

[0174] Utilize the prefetch instructions of ARM (such as PLD) to prefetch the memory data that may be accessed in advance into the cache. According to the execution flow of the program and branch prediction information, predict future memory accesses and prefetch data to reduce memory access latency. The following is an example of assembly code for prefetch optimization:

[0175] / / Assume that the predicted next access address is stored in r10

[0176] Prefetch Load Data [r10]

[0177] / / Subsequent execution code

[0178] Load Register r11, [r10]

[0179] Existing instruction conversions often only focus on the conversion of the basic functions of instructions. However, the present invention selects the best instruction combination according to the characteristics of the ARM architecture through instruction selection optimization, improving the parallelism of memory access and data transfer efficiency. At the same time, the introduction of prefetch optimization enables active reduction of memory access latency. Utilizing the prefetch function of the ARM core, the overall performance is improved.

[0180] In summary, the above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A dynamic instruction conversion memory access optimization method based on data pre-alignment, characterized in that: The specific steps include: Step 1, construct a memory mapping table from the x86 segment base address to the ARM page table base address in the ARM system, load and execute the executable file through dynamic instruction conversion, form a first variable table by initialized variables in global variables and static variables, and form a second variable table by uninitialized variables; establish a data information table to save the instruction name and its ARM starting address, data length and aligned ARM target address, as well as the variable name and its ARM starting address, data length and aligned ARM target address; Step 2: parse the current instruction to be converted to obtain the x86 segment base address and the offset address within the x86 segment corresponding to the instruction and operand. If there is a corresponding entry in the memory mapping table, obtain the ARM page table base address; otherwise, assign the ARM page table base address to the x86 segment base address and add it to the memory mapping table; Step 3, calculate the ARM starting address of the instruction and variable according to the ARM page table base address and the offset address in the x86 segment, and check whether there is a corresponding table entry in the data information table according to the instruction name and its ARM starting address, the variable name and its ARM starting address. If it exists, and when the table entry contains the ARM target address, execute step 4, if not, execute step 5; otherwise, execute step 6; Step 4: When the variable belongs to the first variable table, the instruction and the variable alignment operation are executed according to the ARM target address, and the initial value of the variable is copied to the aligned buffer in sequence; When belonging to the second variable table, an instruction alignment operation is performed; the current instruction to be converted is converted into an ARM instruction with the same function and the ARM target address as an operand; Step 5: According to the ARM start addresses of the instructions and variables in the current instruction to be converted, the alignment addresses of the instructions and variables are calculated as the ARM target addresses, and the ARM target addresses are saved in the data information table, and the current instruction to be converted is converted into an ARM instruction with the same function using the ARM target address as an operand; Step 6: Construct a first ARM instruction for checking whether the variable address is aligned, a second ARM instruction with the result of the first ARM instruction as a jump condition, and a third ARM instruction for address alignment, and convert the current instruction to be converted into an ARM instruction sequence consisting of an ARM instruction with the same function, the first ARM instruction, the second ARM instruction, and the third ARM instruction.

2. The dynamic instruction conversion memory access optimization method according to claim 1, characterized in that: The memory mapping table is stored in a memory space with lower access latency in the ARM system.

3. The dynamic instruction conversion memory access optimization method according to claim 1, characterized in that: The table entries of the memory mapping table are represented as 64-bit data, wherein the lower 32 bits store the ARM page table base address and the upper 32 bits store the x86 segment base address.

4. The dynamic instruction conversion memory access optimization method according to claim 1, characterized in that: Also includes: The instructions in the backward setting range of the converted ARM instructions or ARM instruction sequences are taken as the instruction optimization area. If there are multiple memory data transfer instructions with continuous memory space as operands in the instruction optimization area, these memory data transfer instructions are merged into instructions for batch data transfer to the continuous memory space; otherwise, the instructions in the instruction optimization area are kept unchanged.

5. The dynamic instruction conversion memory access optimization method according to claim 1, characterized in that: Also includes: A data prefetch instruction is added before the converted ARM instruction or ARM instruction sequence.

6. The dynamic instruction conversion memory access optimization method according to claim 1, characterized in that: Also includes: When the ARM system captures a data abort exception and the exception address corresponds to data in the first variable table or the second variable table, the ARM start address and data length in the exception information are saved in the data information table, and then the execution of the program ends.

7. The dynamic instruction conversion memory access optimization method according to claim 1, characterized in that: When the executable file has a complete symbol table, the data information table is built before the executable file is loaded and executed.

8. The dynamic instruction conversion memory access optimization method according to claim 2, characterized in that: The memory space with lower access latency is obtained as follows: Step 1.1, initialize the SysTick timer to generate an interrupt every 1ms; Step 1.2, perform multiple memory read and write operations on the addresses in the memory space under test with the set value as the step length, calculate the number of clock cycles consumed by the operation by recording the value of the SysTick timer before and after the operation, and obtain the access time; Step 1.3: The address with the shortest access time is the starting address of the memory space with lower access latency.

9. The dynamic instruction conversion memory access optimization method according to claim 8, characterized in that: The memory space with lower access latency is a tightly coupled memory.

Citation Information

Patent Citations

  • Cross-memory page difference compatible operation method based on memory access instruction reconstruction

    CN119248286A

  • Prefetch instruction conversion optimization method based on memory access mode virtualization

    CN119440626A