An Optimization Method for Memory Access of ARM Many-Core Instruction Conversion Based on Reinforcement Learning

By applying reinforcement learning-based instruction conversion method in ARM multi-core systems, the problem of difficulty in making full use of multi-core and cache architecture in the prior art is solved, and higher execution performance and memory access efficiency are achieved.

CN119690517BActive Publication Date: 2025-05-30北京麟卓信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510206901.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-30
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

When executing x86 programs in ARM multi-core systems, existing dynamic instruction conversion methods are difficult to make full use of the advantages of multi-core and cache architecture, resulting in a degradation in execution performance.

Method used

Using the ARM multi-core instruction conversion memory access optimization method based on reinforcement learning, by establishing a multi-core instruction pre-allocation model, calculating the calculation core required for the instructions to be converted according to the ARM system state, and adjusting the instruction pre-conversion results, and constructing pre-fetch instructions to optimize memory access.

Benefits of technology

It reduces memory access latency and improves the overall performance of the system, especially for memory-intensive programs, it effectively utilizes the multi-core and cache advantages of ARM multi-core systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119690517B_ABST
    Figure CN119690517B_ABST
Patent Text Reader

Abstract

The present invention discloses an ARM multi-core instruction conversion memory access optimization method based on reinforcement learning. By constructing a multi-core instruction pre-allocation model based on reinforcement learning, when loading and executing an executable file in the way of dynamic instruction conversion, the instruction to be converted is first pre-converted into a first ARM instruction. At the same time, based on the multi-core instruction pre-allocation model, the computing core determined for the instruction to be converted according to the ARM system state is used to adjust the pre-conversion result of the instruction to be converted. Then, a prefetch instruction is constructed according to the virtual address of the instruction to be converted, the write position of the prefetch instruction is determined, and the instruction to be converted is converted into an instruction sequence composed of the prefetch instruction and the ARM instruction, reducing the memory access latency and improving the overall performance of the system. Especially for memory-intensive programs, it can better utilize the multi-core and cache advantages of the ARM multi-core system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer software development, and particularly relates to a memory access optimization method for ARM multi-core instruction conversion based on reinforcement learning. Background Art

[0002] With the growth of heterogeneous computing requirements, using an ARM multi-core system to execute traditional x86 programs has become an important way to expand computing power. An ARM multi-core system usually refers to a system with more than 64 computing cores. However, when executing an x86 program in an ARM multi-core system in a dynamic instruction conversion manner, x86 and ARM have different instruction set architectures, and memory access problems will be faced during the instruction conversion process. Especially when executing the conversion in an ARM multi-core (more than 64 cores) system, the existing conversion methods are difficult to make full use of its powerful multi-core and cache architecture advantages, which will lead to a decline in the execution performance when the program executes across architectures. Summary of the Invention

[0003] In view of this, the present invention provides a memory access optimization method for ARM multi-core instruction conversion based on reinforcement learning, which realizes the conversion of x86 memory access instructions based on computing core scheduling by using reinforcement learning.

[0004] A memory access optimization method for ARM multi-core instruction conversion based on reinforcement learning provided by the present invention specifically includes the following steps:

[0005] Step 1: Establish a multi-core instruction pre-allocation model based on reinforcement learning. Its environment is an ARM multi-core system, the state is the utilization rate of computing cores, the action is to allocate instructions to the selected computing cores, and the reward function is a reward function based on the utilization rate of computing cores, and complete the training of the multi-core instruction pre-allocation model; in the ARM multi-core system, load and execute an executable file through dynamic instruction conversion, and extract the currently to-be-converted instruction of the memory access instruction type.

[0006] Step 2: Use the multi-core instruction pre-allocation model to calculate the first computing core to which the currently to-be-converted instruction needs to be allocated according to the current state of the ARM multi-core system, and its corresponding first address range.

[0007] Step 3: Pre-convert the currently to-be-converted instruction into a first ARM instruction to obtain its first memory address. If the first ARM instruction is allocated to the first computing core, execute Step 4; otherwise, allocate the first ARM instruction to the first computing core and then execute Step 4.

[0008] Step 4: Construct a first prefetch instruction from the first memory address and the prefetch distance. The memory address of the first prefetch instruction is set to the second memory address, and the second memory address includes a first virtual page number and a second page offset within a page, where the second page offset within a page is less than the first page offset within a page; convert the current instruction to be translated into an instruction sequence formed by the first prefetch instruction and the first ARM instruction.

[0009] Further, the calculation method of the reward function in the said Step 1 is: , where is the average utilization rate of the computing cores, U i is the utilization rate of the i-th computing core, and n is the total number of computing cores. is the standard deviation of the utilization rates, is the coefficient for adjusting the influence degree of the standard deviation.

[0010] Further, the state of the many-core instruction pre-allocation model includes the ARM many-core system state and the ARM instruction characteristics. Among them, the ARM many-core system state includes the cache hit rate, the utilization rate of the computing cores, and the instruction queue length of the computing cores, and the ARM instruction characteristics include the memory access range and the computing resource requirements.

[0011] Further, the calculation method of the reward function is: , where w 1 , w 2 , w 3 , w 4 and w 5 are all weight coefficients and w 1 + w 2 + w 3 + w 4 + w 5 = 1, R 1 is the cache hit rate reward, R 2 is the utilization rate reward of the computing cores, R 3 is the instruction queue length reward of the computing cores, R 4 is the memory access range reward, and R 5 is the computing resource requirement reward.

[0012] Further, the calculation method of the cache hit rate reward is: , where H is the current cache hit rate, and H max is the theoretical maximum cache hit rate.

[0013] Further, the calculation method of the instruction queue length reward of the computing cores is: , where Q i is the instruction queue length of the i-th computing core, and Q max is the maximum allowable length of the instruction queue.

[0014] Further, the calculation method of the computing resource requirement reward is as follows: , where C 1 is the computing resource requirement of the instruction, and C 2 is the available computing resource of the allocated computing core.

[0015] Further, the method of allocating the first ARM instruction to the first computing core in step 3 is as follows: Obtain the first virtual page in the ARM virtual memory where the first ARM instruction is located, record the first virtual page number and the first offset within the page. The first virtual page corresponds to the first physical page; Determine the second physical page within the first address range, and modify the page table entry related to the first virtual page in the page table to the mapping relationship from the first virtual page to the second physical page, copy the data of the first physical page to the second physical page, and then clear the first physical page.

[0016] Further, the prefetch distance in step 4 is determined according to the cache hierarchy of the ARM architecture and is set to a larger value when the cache hit rate of the ARM many-core system is low.

[0017] Further, in step 4, the difference between the second offset within the page and the first offset within the page is increased.

[0018] Beneficial effects:

[0019] In the present invention, by constructing a many-core instruction pre-allocation model based on reinforcement learning, when loading and executing an executable file in the way of dynamic instruction conversion, the instruction to be converted is first pre-converted into the first ARM instruction. At the same time, based on the many-core instruction pre-allocation model, the computing core determined according to the ARM system state adjusts the pre-conversion result of the instruction to be converted, and then constructs a prefetch instruction according to the virtual address of the instruction to be converted, determines the write position of the prefetch instruction, and converts the instruction to be converted into an instruction sequence composed of the prefetch instruction and the ARM instruction, reducing the memory access latency and improving the overall performance of the system. Especially for memory-intensive programs, it can better utilize the multi-core and cache advantages of the ARM many-core system. Brief Description of the Drawings

[0020] Figure 1 It is a flow schematic diagram of an ARM many-core instruction conversion memory access optimization method provided by the present invention. Detailed Description of the Invention

[0021] The following lists embodiments in conjunction with the accompanying drawings to describe the present invention in detail.

[0022] An ARM multi-core instruction conversion memory access optimization method based on reinforcement learning provided by the present invention has a core idea as follows: constructing a multi-core instruction pre-allocation model based on reinforcement learning, pre-converting the instruction to be converted into a first ARM instruction when loading and executing an executable file in a dynamic instruction conversion manner, and at the same time adjusting the pre-conversion result of the instruction to be converted according to the computing core determined by the multi-core instruction pre-allocation model based on the ARM system state, then constructing a prefetch instruction according to the virtual address of the instruction to be converted, determining the write position of the prefetch instruction, and converting the instruction to be converted into an instruction sequence composed of the prefetch instruction and the ARM instruction.

[0023] An ARM multi-core instruction conversion memory access optimization method based on reinforcement learning provided by the present invention has a specific process as Figure 1 shown, specifically including the following steps:

[0024] Step 1: Establish a multi-core instruction pre-allocation model based on reinforcement learning for pre-determining the ARM computing core to which the ARM instruction will be allocated; the environment of the multi-core instruction pre-allocation model is the ARM system, the state is the utilization rate of the ARM system computing core, the action is to allocate the ARM instruction to the selected ARM computing core, and the reward function is a reward function based on the computing core utilization rate; complete the training of the multi-core instruction pre-allocation model by using the training method of the reinforcement learning model. When loading and executing an executable file through dynamic instruction conversion in the ARM system, initialize the page table mapping relationship.

[0025] A reasonable instruction scheduling should keep the utilization rates of each computing core balanced, avoiding the situation that some cores are idle while some cores are overloaded, thereby improving the overall utilization rate of the computing core. Therefore, the present invention designs a reward function R according to the utilization rate of the computing core, and the calculation method is: , where is the average utilization rate of the computing core, U i is the utilization rate of the i-th computing core, n is the total number of computing cores, is the standard deviation of the utilization rate, is the coefficient for adjusting the influence degree of the standard deviation.

[0026] In order to further improve the accuracy of instruction allocation, the present invention uses both the ARM system state and the ARM instruction characteristics as the state of the multi-core instruction pre-allocation model. Among them, the ARM system state includes cache hit rate, utilization rate of the computing core, instruction queue length of the computing core, etc., and the ARM instruction characteristics include parameters such as memory access range and computing resource requirements.

[0027] Thus, the reward function of the multi-core instruction pre-allocation model is: , where w 1 、w 2 、w 3, w 4 and w 5 are both weight coefficients, and w 1 + w 2 + w 3 + w 4 + w 5 = 1, R 1 is the cache hit rate reward, R 2 is the utilization rate reward of the computing core, R 3 is the instruction queue length reward of the computing core, R 4 is the memory access range reward, R 5 is the computing resource requirement reward.

[0028] Specifically, the higher the cache hit rate, the better the system performance. The calculation method of the cache hit rate reward is: , where H is the current cache hit rate, and H max is the theoretical maximum cache hit rate.

[0029] The instruction queue length should be kept within a reasonable range. Excessive length will cause congestion. The calculation method of the instruction queue length reward is: , where Q i is the instruction queue length of the i-th computing core, and Q max is the maximum allowable length of the instruction queue.

[0030] The calculation method of the memory access range reward is: R 4 = M, where M is the overlapping ratio of the actual memory access range of the instruction to the cache coverage range.

[0031] The calculation method of the computing resource requirement reward is: , where C 1 is the computing resource requirement of the instruction, and C 2 is the available computing resource of the allocated computing core.

[0032] The page table mapping relationship refers to the process of mapping virtual addresses to physical addresses. In a computer system, in order to implement virtual memory, the operating system uses a page table to record the mapping relationship between virtual addresses and physical addresses.

[0033] Step 2: Obtain the current instruction to be converted. If the current instruction to be converted is a memory access instruction, execute Step 3; otherwise, convert the current instruction to an ARM instruction and then execute Step 7.

[0034] Step 3: Obtain the current ARM system state, and use the multi-core instruction pre-allocation model trained in Step 1 to calculate the computing core to which the current instruction to be converted needs to be allocated according to the current ARM system state. Denote this computing core as the first computing core, and denote the memory address range corresponding to the first computing core as the first address range.

[0035] Step 4: Pre-convert the current instruction to be converted into a first ARM instruction to obtain its first memory address and the first virtual page in the ARM virtual memory where it is located. Record the first virtual page number and the first offset within the page of the first virtual page, and denote the physical page corresponding to the first virtual page as the first physical page;

[0036] If the first physical page is not within the first address range, determine a second physical page within the first address range, modify the page table entry related to the first virtual page in the page table to the mapping relationship from the first virtual page to the second physical page, copy the data of the first physical page to the second physical page, and then clear the first physical page; otherwise, keep the mapping relationship in the page table unchanged.

[0037] Step 5: Determine the prefetch distance for data prefetch according to the cache hierarchy of the ARM architecture. If the cache hit rate of the current ARM system is low, increase the prefetch distance and execute Step 6; otherwise, keep the prefetch distance unchanged and execute Step 6.

[0038] Step 6: Construct a first prefetch instruction from the first memory address and the prefetch distance, and set the memory address of the first prefetch instruction to a second memory address. The second memory address includes the first virtual page number and a second offset within the page, and the second offset within the page is less than the first offset within the page; convert the current instruction to be converted into an instruction sequence formed by the first prefetch instruction and the first ARM instruction.

[0039] In addition, to improve the cache hit rate, the difference between the second offset within the page and the first offset within the page can be further increased to realize loading the memory data to be processed into the cache earlier.

[0040] Step 7: If the conversion execution of the executable file is completed, end this process; otherwise, execute Step 2.

[0041] Embodiment:

[0042] In this embodiment, a method for optimizing memory access in ARM multi-core instruction conversion based on reinforcement learning provided by the present invention is adopted to achieve the efficient execution of x86 architecture executable files on an ARM multi-core system, including the following steps:

[0043] S1. Instruction analysis and marking: Deeply analyze and mark the memory access instructions and operations of the input x86 assembly code. The specific steps are as follows:

[0044] The following is an assembly code containing typical x86 memory access instructions, for example:

[0045] section.data

[0046] array db 1, 2, 3, 4, 5, 6, 7, 8, 9, 10

[0047] section.text

[0048] global_start

[0049] _start:

[0050] mov eax, [array] (change to address f)

[0051] add eax, 1

[0052] mov [array], eax

[0053] mov ebx, [array + 4]

[0054] sub ebx, 2

[0055] mov [array + 8], ebx

[0056] jmp loop_start

[0057] loop_start:

[0058] mov ecx, 10

[0059] loop_iteration:

[0060] mov edx, [array + ecx]

[0061] add edx, 3

[0062] mov [array + ecx], edx

[0063] loop loop_iteration

[0064] Parse the input x86 instructions to identify memory access instructions such as mov eax, [array], mov [array], eax, etc., and mark the memory addresses they access and the operation types.

[0065] For the above code, through parsing, it can be determined that mov eax, [array] and mov edx, [array + ecx] are read operations, while mov [array], eax and mov [array + 8], ebx are write operations.

[0066] S2, Address remapping, which maps x86 memory addresses to the ARM memory space and dynamically adjusts according to the ARM multi-core system. The specific steps are as follows:

[0067] S2.1, Map the x86 memory addresses to the ARM memory space, taking into account the cache hierarchy and core distribution of the ARM multi-core architecture. Assume that the ARM multi-core system has 128 cores, divided into 16 groups, with 8 cores in each group. Each core group has its own L1 and L2 caches and shares the L3 cache.

[0068] S2.2, The dynamic address mapping algorithm includes:

[0069] S2.2.1, Initialize the mapping table.

[0070] Create a mapping table to store the mapping information from x86 addresses to ARM addresses. For the initial mapping, divide the entire x86 memory space into multiple regions, with each region corresponding to an ARM core group. For example, map the low-address region of x86 to the memory space of core group 0, map the slightly higher-address region to the memory space of core group 1, and so on.

[0071] For the starting address of the array, assume it is in the low-address region and map it to the L2 cache associated region of core group 0, initially mapped to the ARM address 0x80000000.

[0072] S2.2.2, Core group load monitoring. Use hardware performance counters to monitor the load conditions of each core group, including cache hit rate, memory bandwidth usage, and instruction execution rate, etc. For example, continuously obtain the L2 cache hit rate H0, memory bandwidth B0, and instruction execution rate I0 of core group 0 through hardware performance counters.

[0073] S2.2.3, Dynamic mapping adjustment

[0074] During program execution, dynamically adjust the mapping according to the load conditions of the core groups. If the load of core group 0 is too high, for example, the cache hit rate H0 is lower than a certain threshold TH1, or the memory bandwidth B0 exceeds a certain threshold TH2, migrate some address mappings to the core group with lighter load.

[0075] For addresses such as array+4, array+8, etc., calculate their offsets relative to array. When the mapping of array needs to be adjusted, map them to the corresponding positions of the new core group according to the offsets. For example, if array migrates from core group 0 to core group 1, array+4 will be mapped to the corresponding position of core group 1 according to its offset, taking into account cache line alignment. Assume the cache line size is 64 bytes and ensure that the mapped address is aligned to the cache line boundary.

[0076] For dynamic addresses, such as array + ecx in loop_iteration, the mapped core group is dynamically selected according to the value of ecx and the current load of the core group. In each iteration, the load of the core group is checked. If the load of core group 0 is high while the load of core group 2 is low, array + ecx is mapped to the corresponding position in core group 2, and the mapping table is updated accordingly.

[0077] Using the dynamic address mapping algorithm, the mapping address is adjusted in real time according to the load of the core group and the cache status. Frequently accessed addresses are mapped to the core group with lighter load and higher cache hit rate to improve cache utilization.

[0078] S3. Instruction conversion: Convert x86 instructions to ARM instructions and add prefetch instructions. The specific steps are as follows:

[0079] The result of converting the example x86 instruction in step one to an ARM instruction is as follows:

[0080] .data

[0081] array:.byte 1, 2, 3, 4, 5, 6, 7, 8, 9, 10

[0082] .text

[0083] .global_start

[0084] _start:

[0085] ldr r0, [r1]; corresponding to move ax, [array]

[0086] add r0, #1

[0087] str r0, [r1]; corresponding to mov [array], eax

[0088] ldr r2, [r1, #4]; corresponding to move bx, [array + 4]

[0089] sub r2, #2

[0090] str r2, [r1, #8]; corresponding to mov [array + 8], ebx

[0091] b loop_start_arm

[0092] loop_start_arm:

[0093] mov r3, #10

[0094] loop_iteration_arm:

[0095] ldr r4, [r1, r3]; corresponding to mov edx, [array + ecx]

[0096] add r4, #3

[0097] str r4, [r1, r3]; corresponding to mov [array + ecx], edx

[0098] sub r3, r3, #1

[0099] bne loop_iteration_arm

[0100] During the conversion process, the conversion is carried out according to the characteristics of the ARM instruction set, and the memory access characteristics are also considered. The detailed steps for adding prefetch instructions according to the address remapping result are as follows:

[0101] S3.1. Address analysis. When converting instructions, for each memory access instruction, query the mapping table generated in the address remapping stage to obtain its mapped address in the ARM architecture. For example, for ldr r0, [r1] corresponding to mov eax, [array], look up the ARM mapped address 0x80000000 of array.

[0102] S3.2. Prefetch distance determination. Determine the prefetch distance according to the cache hierarchy and performance data of the ARM architecture. For example, if the prefetch unit of the L2 cache can prefetch data with a distance of 128 bytes, for ldr r0, [r1], add a prefetch instruction to prefetch the data near [r1].

[0103] S3.3. Prefetch instruction insertion. Insert prefetch instructions before the corresponding memory access instructions. For ldr r0, [r1], insert pld [r1, #0] to prefetch the data at [r1]. For ldr r2, [r1, #4], insert pld [r1, #4] to prefetch the data at [r1 + 4]. For ldr r4, [r1, r3] in the loop, insert pld [r1, r3] at the beginning of each iteration to prefetch the data at [r1 + r3].

[0104] S3.4. Dynamic prefetch adjustment. Combine hardware performance counters, such as cache miss rate and memory latency data, to dynamically adjust the prefetch distance and timing. If it is found that the cache miss rate is high, increase the prefetch distance or advance the insertion position of the prefetch instruction. For example, if the cache miss rate caused by pld [r1, #4] is still high, adjust it to pld [r1, #8] or insert the prefetch instruction at an earlier position.

[0105] For memory access instructions, prefetch instructions are added according to the address remapping result.

[0106] S4. Dynamically allocate instruction blocks to the core groups of the ARM multi-core according to performance metrics, and adopt a core allocation method based on reinforcement learning. The specific steps are as follows:

[0107] S4.2. Model initialization.

[0108] State representation. Define the system state, including the load state of the core group and the characteristics of the instruction block. For example, the state S can be represented as:

[0109] {core_group_0:{H0,B0,Q0},core_group_1:{H1,B1,Q1},instruction_block:{memory_range,resource_demand}}.

[0110] Action space definition. Define the action space, that is, the operation of allocating the instruction block to different core groups. For example, the action A can be to allocate the instruction block to core group 0, core group 1, etc.

[0111] Reward function design. Design the reward function and determine the reward according to the system performance metrics. If the system performance improves after the instruction block is allocated to the core group, a positive reward is given; otherwise, a negative reward is given. For example, the reward function R can be calculated based on the change in cache hit rate and the change in execution time.

[0112] Reinforcement learning training. Use reinforcement learning algorithms, such as Q-learning or deep reinforcement learning algorithms, for training. The agent selects the action A according to the current state S, observes the reward function R and the new state S' after executing the action, and updates the policy. For example, in the initial state, the instruction block is allocated to core group 0 according to the policy, the change in performance metrics is observed, and the policy is updated according to the reward function, so that the agent learns the optimal core allocation policy.

[0113] S4.2. Dynamic adjustment.

[0114] During the program execution, continuously use the reinforcement learning agent for the core allocation policy. Update the state S according to the real-time performance data, and select the action A according to the learned policy to dynamically adjust the allocation of the instruction block.

[0115] For the optimized instruction blocks mentioned above, evaluate their memory access range and frequency. For example, for the loop_iteration_arm instruction block, it mainly accesses the array and its nearby addresses, and according to the address remapping information, it is allocated to core group 0.

[0116] Continuously monitor the performance metrics of the core group, such as cache hit rate, memory bandwidth usage, and instruction execution latency. If there is a performance bottleneck in core group 0, migrate some instruction blocks to other core groups.

[0117] Adopt a core allocation strategy based on reinforcement learning. According to the performance feedback of the core group and task characteristics, dynamically adjust the allocation of instruction blocks to achieve the optimal utilization of core resources. Utilize hardware performance counters to collect performance data in real time, providing a basis for core allocation decisions and improving the adaptability and performance of the system.

[0118] In summary, the above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for optimizing ARM multi-core instruction conversion memory access based on reinforcement learning, characterized in that: The specific steps include: Step 1: Establish a many-core instruction pre-allocation model based on reinforcement learning, where the environment is an ARM many-core system, the state is the utilization of the computing core, the action is to allocate instructions to the selected computing core, and the reward function is a reward function based on the utilization of the computing core. The training of the many-core instruction pre-allocation model is completed; load and execute the executable file in the ARM many-core system through dynamic instruction conversion, and extract the current instructions to be converted of the memory access instruction type; Step 2: Use the many-core instruction pre-allocation model to calculate the first computing core to which the current instruction to be converted needs to be allocated and the first address range corresponding to it according to the current ARM many-core system state; Step 3: pre-convert the current instruction to be converted into a first ARM instruction to obtain its first memory address and the first virtual page in the ARM virtual memory where it is located, record the first virtual page number and the first page offset of the first virtual page, and record the physical page corresponding to the first virtual page as the first physical page; If the first physical page is not within the first address range, determine the second physical page in the first address range, modify the page table entry related to the first virtual page in the page table to a mapping relationship from the first virtual page to the second physical page, copy the data of the first physical page to the second physical page, and then clear the first physical page; otherwise, keep the mapping relationship in the page table unchanged; Determine the prefetch distance of data prefetch according to the cache hierarchy of the ARM architecture; Step 4: construct a first prefetch instruction based on the first memory address and the prefetch distance, wherein the memory address of the first prefetch instruction is set to a second memory address, wherein the second memory address includes a first virtual page number and a second page offset, and the second page offset is smaller than the first page offset; The current instruction to be converted is converted into an instruction sequence formed by a first prefetch instruction and a first ARM instruction.

2. The ARM multi-core instruction conversion memory access optimization method according to claim 1, characterized in that: The calculation method of the reward function in step 1 is: in, To calculate the average utilization of the core, U i is the utilization of the i-th computing core, n is the total number of computing cores, is the standard deviation of utilization, and α is the coefficient for adjusting the influence of standard deviation.

3. The ARM multi-core instruction conversion memory access optimization method according to claim 1, characterized in that: The state of the many-core instruction pre-allocation model includes the ARM many-core system state and ARM instruction characteristics, wherein the ARM many-core system state includes cache hit rate, computing core utilization and computing core instruction queue length, and the ARM instruction characteristics include memory access range and computing resource requirements.

4. The ARM multi-core instruction conversion memory access optimization method according to claim 3, characterized in that: The reward function is calculated as follows: R=w1R1+w2R2+w3R3+w4R4+w5R5, wherein w1, w2, w3, w4 and w5 are all weight coefficients and w1+w2+w3+w4+w5=1, R1 is the cache hit rate reward, R2 is the computing core utilization reward, R3 is the computing core instruction queue length reward, R4 is the memory access range reward, and R5 is the computing resource demand reward.

5. The ARM multi-core instruction conversion memory access optimization method according to claim 4, characterized in that: The calculation method of the cache hit rate reward is: Among them, H is the current cache hit rate, H max is the theoretical maximum cache hit rate.

6. The ARM multi-core instruction conversion memory access optimization method according to claim 4, characterized in that: The calculation method of the instruction queue length reward of the computing core is: Among them, Q i is the instruction queue length of the i-th computing core, Q max The maximum allowed length of the instruction queue.

7. The ARM multi-core instruction conversion memory access optimization method according to claim 4, characterized in that: The calculation method of the computing resource demand reward is: Among them, C1 is the computing resource requirement of the instruction, and C2 is the available computing resources of the allocated computing core.

8. The ARM multi-core instruction conversion memory access optimization method according to claim 1, characterized in that: The prefetch distance in step 4 is determined according to the cache hierarchy of the ARM architecture, and the prefetch distance is increased when the cache hit rate of the ARM many-core system is low.

Citation Information

Patent Citations

  • Multi-core DSP task scheduling method and device based on reinforcement learning

    CN118331696A

  • Prefetch instruction conversion optimization method based on memory access mode virtualization

    CN119440626A