An Instruction Conversion Memory Access Conflict Optimization Method Based on Instruction Stream Feature Recognition

By using instruction conversion method based on instruction flow feature recognition in ARM multi-core system, the memory access instructions related to shared memory is identified and processed, the problem of memory access conflicts in ARM multi-core system is solved, and the system performance is significantly improved.

CN119883435BActive Publication Date: 2025-06-20北京麟卓信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510377539.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-20
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

When the ARM multi-core system executes x86 programs, multiple computing cores access memory at the same time, memory access conflicts are prone to occur, resulting in performance degradation. The existing memory access conflict resolution methods are difficult to effectively deal with complex memory access situations, resulting in a significant increase in memory access latency.

Method used

The instruction conversion memory access conflict optimization method based on instruction flow feature recognition is adopted. The instruction flow analysis determines the instruction sequence related to the x86 system call and converts it into an ARM instruction sequence to identify the memory-related memory addresses to form a shared memory list. For the memory fetch instruction located in the shared memory list, it is converted into an instruction sequence composed of a lock instruction sequence, a second ARM instruction sequence and an unlock instruction sequence, and the conversion execution is completed based on the cache consistency protocol.

Benefits of technology

It significantly reduces memory access conflicts of x86 programs on ARM multi-core systems, improves the performance of dynamic instruction conversion, and ensures high-performance operation of ARM multi-core systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883435B_ABST
    Figure CN119883435B_ABST
Patent Text Reader

Abstract

The present invention discloses an instruction conversion memory access conflict optimization method based on instruction stream feature recognition. During the pre-execution of an executable file in an ARM many-core system in a dynamic instruction conversion manner, instruction stream analysis is used to determine the shared memory among processes to form a first shared memory list. On this basis, the executable file is loaded and executed again. For the memory access type instructions to be converted whose memory addresses are in the first shared memory list, they are converted into an instruction sequence composed of a locking instruction sequence, a second ARM instruction sequence, and an unlocking instruction sequence. At the same time, the conversion execution of the executable file is completed based on the cache coherence protocol, significantly reducing the memory access conflicts of x86 programs on the ARM many-core system and providing strong support for achieving high-performance dynamic instruction conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer instruction conversion and multi-core systems, and particularly relates to an instruction conversion memory access conflict optimization method based on instruction stream feature recognition. Background Art

[0002] When an ARM multi-core system executes an x86 program in a dynamic instruction conversion manner, multiple ARM computing cores will access memory simultaneously, and at this time, memory access conflicts are likely to occur, thereby degrading the performance of the ARM multi-core system. Existing memory access conflict resolution methods are difficult to effectively handle complex memory access situations, which will lead to a significant increase in memory access latency, and the overall performance of the ARM multi-core system cannot meet expectations. Summary of the Invention

[0003] In view of this, the present invention provides an instruction conversion memory access conflict optimization method based on instruction stream feature recognition, which realizes the feature recognition of shared memory through instruction stream analysis and completes the dynamic conversion of x86 instructions to ARM instructions.

[0004] An instruction conversion memory access conflict optimization method based on instruction stream feature recognition provided by the present invention specifically includes the following steps:

[0005] Step 1: Pre-execute an executable file in an ARM multi-core system in a dynamic instruction conversion manner, use instruction stream analysis to determine a first x86 instruction sequence related to x86 system calls, and parse the first ARM instruction sequence obtained by converting the first x86 instruction sequence to determine the shared memory among processes to form a first shared memory list;

[0006] Step 2: Load and execute the executable file in a dynamic instruction conversion manner. If the current instruction to be converted is a memory access instruction and the memory address is in the first shared memory list, then convert the current instruction to be converted into an instruction sequence composed of a locking instruction sequence, a second ARM instruction sequence, and an unlocking instruction sequence, and execute Step 3 when the current instruction to be converted is a read operation and execute Step 4 when it is a write operation; otherwise, execute Step 5;

[0007] Step 3: If the cache behavior corresponding to the current computing core is in an invalid state, send a read request to the bus to read data and set the state of the cache line to the shared state or the exclusive state after the reading is completed; if it is in the shared state or the exclusive state, read data from the cache line; then execute Step 5;

[0008] Step 4: If the cache line corresponding to the current computing core is in an exclusive state, the data in the cache line is modified and set to a modified state; if it is in a shared state, an invalidation message is sent to other computing cores, and after the response is confirmed, it is set to a modified state, and then a write operation is performed; if it is in an invalid state, a read-exclusive request is sent to the bus, and after obtaining the data, it is set to a modified state, and then a write operation is performed; if it is in a modified state, a write operation is performed;

[0009] Step 5: If the executable file has not completed execution, execute step 2; otherwise, end this process.

[0010] Furthermore, the method of using instruction stream analysis to determine the first x86 instruction sequence related to the x86 system call in step 1 is: after using instruction stream analysis to identify the system call number of the shared memory related x86 system call, the first x86 instruction sequence is formed from the instruction containing the system call number to the soft interrupt trigger instruction.

[0011] Furthermore, the manner in which the first ARM instruction sequence converted from the first x86 instruction sequence is parsed in step 1 to determine the shared memory between processes to form the first shared memory list is:

[0012] If the first ARM instruction sequence includes a system call for creating or obtaining a shared memory segment, an ARM instruction A for saving a shared memory key value is added after the first ARM instruction sequence; if the first ARM instruction sequence includes a system call for attaching a shared memory segment to an address space of a calling process, an ARM instruction B for saving a shared memory address is added after the first ARM instruction sequence; the process ID, shared memory key value and shared memory address of the current process are obtained, a shared memory list is formed by the shared memory key value, the shared memory address and the process ID for creating the shared memory, and the shared memory list is used as the first shared memory list.

[0013] Furthermore, the process ID in the shared memory list is traversed as the target process, and the page table entries of all shared memory pages of the target process are obtained and recorded as original page table entries; the mapping relationship between the above shared memory page and the physical memory page is released, and then the information of the original page table entry is saved in the page table of the target process, and the shared memory address that triggers the page fault exception and the exception address belongs to the shared memory list is saved in the first shared memory list.

[0014] Furthermore, the locking instruction sequence in step 2 is composed of an instruction for setting a lock variable, an atomic operation LDREXB for loading an exclusive byte, and an atomic operation STREXB for storing an exclusive byte.

[0015] Furthermore, the unlocking instruction sequence in step 2 is composed of instructions for modifying lock variables.

[0016] Further, the second ARM instruction sequence in step 2 is composed of ARM instructions with the same function as the currently to-be-converted instruction and other memory access-related instructions.

[0017] Further, the method of sending a read request to the bus to read data in step 3 is to read data from the cache line or memory of other computing cores.

[0018] Further, the x86 system calls in step 1 include: shmget, shmat, shmdt, and shmctl. Beneficial effects

[0019] In the process of pre-executing an executable file in an ARM multi-core system in a dynamic instruction conversion manner, the present invention uses instruction stream analysis to determine the shared memory between processes to form a first shared memory list. On this basis, the executable file is loaded and executed again. For the memory access type to-be-converted instructions whose memory addresses are in the first shared memory list, they are converted into an instruction sequence composed of a locking instruction sequence, a second ARM instruction sequence, and an unlocking instruction sequence. At the same time, the conversion execution of the executable file is completed based on the cache coherence protocol, significantly reducing the memory access conflicts of x86 programs on the ARM multi-core system and providing strong support for achieving high-performance dynamic instruction conversion. Description of the drawings

[0020] Figure 1 It is a schematic flowchart of a method for optimizing instruction conversion memory access conflicts based on instruction stream feature recognition provided by the present invention. Detailed implementation manners

[0021] The following lists embodiments in conjunction with the drawings to describe the present invention in detail.

[0022] A method for optimizing instruction conversion memory access conflicts based on instruction stream feature recognition provided by the present invention, its core idea is: in the process of pre-executing an executable file in an ARM multi-core system in a dynamic instruction conversion manner, using instruction stream analysis to determine the shared memory between processes to form a first shared memory list. On this basis, the executable file is loaded and executed again. For the memory access type to-be-converted instructions whose memory addresses are in the first shared memory list, they are converted into an instruction sequence composed of a locking instruction sequence, a second ARM instruction sequence, and an unlocking instruction sequence. At the same time, the conversion execution of the executable file is completed based on the cache coherence protocol.

[0023] A method for optimizing instruction conversion memory access conflicts based on instruction stream feature recognition provided by the present invention, the specific process is as Figure 1 shown, and specifically includes the following steps:

[0024] Step 1, pre-execute an executable file in an ARM many-core system in a dynamic instruction conversion manner, identify the system call number of the shared memory related x86 system call through instruction stream analysis, form a first x86 instruction sequence from the instruction containing the system call number to the soft interrupt trigger instruction, and convert the first x86 instruction sequence into a first ARM instruction sequence; if the first ARM instruction sequence contains a system call for creating or obtaining a shared memory segment, then add an ARM instruction A for saving a shared memory key value after the first ARM instruction sequence; if the first ARM instruction sequence contains a system call for attaching a shared memory segment to the address space of a calling process, then add an ARM instruction B for saving a shared memory address after the first ARM instruction sequence; obtain the process ID, shared memory key value and shared memory address of the current process, and save the shared memory key value, shared memory address and the process ID for creating the shared memory in a shared memory list.

[0025] There are multiple system calls in the existing x86 architecture for implementing shared memory related operations, including: shmget, shmat, shmdt, shmctl, etc., among which shmget is used to create or obtain a shared memory segment, shmat is used to attach a shared memory segment to the address space of the calling process, shmdt is used to separate a shared memory segment from the address space of the calling process, and shmctl is used to control the shared memory segment. Usually, the system calls related to shared memory in the ARM architecture and the x86 architecture are the same. For example, the system call in the ARM architecture similar to the shmget in the x86 architecture is also shmget.

[0026] Step 2: traverse the process ID in the shared memory list as the target process, obtain the page table entries of all shared memory pages of the target process, and record them as original page table entries; release the mapping relationship between all the above shared memory pages and physical memory pages, so that all processes cannot use all the above shared memory pages, and then save the information of the original page table entries into the page table of the target process.

[0027] Each process has its own independent virtual address space, and the process's page table is used to convert the virtual address used by the process into the actual physical memory address.

[0028] Step 3: When a page fault exception is triggered and the exception address belongs to the shared memory list, the shared memory address corresponding to the exception memory page is saved in the first shared memory list, and then the mapping relationship between the exception memory page and the physical memory page is restored to complete the pre-execution of the executable file.

[0029] In the present invention, by pre-executing an executable file in an ARM many-core system in a dynamic instruction conversion manner, it is possible to pre-obtain an instruction sequence related to system calls in the executable file, that is, a first x86 instruction sequence. The first x86 instruction sequence is converted into a first ARM instruction sequence according to the existing dynamic instruction conversion method. Then, different ARM instructions are added according to the different system calls included in the first ARM instruction sequence to save relevant information of the shared memory, and the relevant shared memory is initially determined. By modifying the page table entries of the shared memory and creating the page table of the process of the shared memory to construct a triggering condition for a page fault exception, it is possible to further determine the shared memory that will actually be shared among multiple processes to form a first shared memory list.

[0030] Step 4: Load and execute the executable file in the ARM many-core system in a dynamic instruction conversion manner, obtain the current instruction to be converted. If the current instruction to be converted is a memory access instruction and the relevant memory address is in the first shared memory list, then when the current instruction to be converted is a read operation, execute Step 5, and when the current instruction to be converted is a write operation, execute Step 6; otherwise, execute Step 7.

[0031] Step 6: Convert the current instruction to be converted into an instruction sequence composed of a locking instruction sequence, a second ARM instruction sequence, and an unlocking instruction sequence. At the same time, if the state of the cache line corresponding to the current instruction to be converted in the current computing core is the invalid state, the current computing core sends a read request to the bus to read the data and sets the state of the cache line where the data is located to the shared state or the exclusive state after the reading is completed; if it is the shared state or the exclusive state, the current computing core directly reads the data from the cache line; then execute Step 7.

[0032] Among them, the locking instruction sequence is composed of an instruction for setting a lock variable, an atomic operation LDREXB for loading an exclusive byte, and an atomic operation STREXB for storing an exclusive byte, and the unlocking instruction sequence is composed of an instruction for modifying the lock variable. The way for the current computing core to send a read request to the bus to read the data is to read the data from the cache line or memory of other computing cores. The second ARM instruction sequence is composed of ARM instructions with the same function as the current instruction to be converted and other memory access-related instructions.

[0033] For example, computing core A executes a read instruction to access data at a certain address, and the state of its cache line is the invalid state (Invalid). Computing core A issues a read request. If the data is in the cache of computing core B and is in the modified state (Modified), computing core B first writes the data back to the main memory, then sends the data to computing core A, and at the same time, the cache line states of both computing core B and computing core A become the shared state (Shared).

[0034] In addition, the cache line of computing core A is in the shared state. When a read instruction is executed to access the data corresponding to this cache line, it is directly read from the cache, and the state remains the shared state.

[0035] Step 6: Convert the current instruction to be converted into an instruction sequence composed of a locking instruction sequence, a second ARM instruction sequence, and an unlocking instruction sequence. At the same time, if the state of the cache line corresponding to the current instruction to be converted in the current computing core is the exclusive state, the current computing core directly modifies the data in the cache line and sets the state of the cache line to the modified state; if it is the shared state, an invalidate message is sent to other computing cores. After all other computing cores respond and confirm, the state of the cache line is set to the modified state, and then a write operation is performed; if it is the invalid state, a read-exclusive request is sent to the bus, and after obtaining the data, the state of the cache line is set to the modified state, and then a write operation is performed; if it is the modified state, a write operation is directly performed.

[0036] Specifically, when the cache line state is the exclusive state (Exclusive), since the data only exists in the cache of the current computing core and is consistent with the main memory data, the computing core can directly modify the data in the cache line and change the cache line state to the modified state. At this time, there is no need to communicate with other computing cores. For example, the cache line of computing core A is in the exclusive state. After executing a write instruction to modify the data in this cache line, the cache line state becomes the modified state.

[0037] When a computing core needs to perform a write operation on a cache line in the shared state, it needs to first send an invalidate message (Invalidate message) to other computing cores to notify other computing cores to invalidate the copy of this data. After receiving the message, other computing cores will set the state of the corresponding cache line in their own cache to the invalid state. When all other computing cores respond and confirm, the computing core that initiates the write operation changes the state of its own cache line to the modified state, and then performs the write operation. For example, there are copies of a certain data in the cache lines of computing cores A, B, and C, and the state is the shared state. When computing core A executes a write instruction, it sends an invalidate message to computing cores B and C. After receiving the message, computing cores B and C set the state of the corresponding cache lines to the invalid state and reply with confirmation. After computing core A receives the confirmation, it changes the state of its own cache line to the modified state and then performs the write operation.

[0038] When the cache line is invalid, the computing core needs to obtain a valid copy of the data first, so it will issue a read-exclusive request to the bus, which means that it wants to read the data and has exclusive access to the subsequent write of the data. If there is a valid copy of the data in the cache line of other computing cores, the computing core with the copy will send the data to the core that initiated the request and set its own cache line state to invalid; if other computing cores do not have a valid copy, the data will be read from the main memory. After obtaining the data, the cache line state is set to modified, and then a write operation is performed. For example, the cache line of computing core A is in an invalid state, and a read-exclusive request is issued when executing a write instruction. If computing core B has a valid copy of the data, computing core B sends the data to computing core A and sets its own cache line to invalid. After receiving the data, computing core A changes the cache line state to modified, and then performs a write operation.

[0039] When the cache line is in the modified state, the computing core can directly modify the data in the cache line because the computing core has the only valid copy of the data and it has been modified. The cache line state remains in the modified state. For example, the cache line of computing core A is in the modified state. When executing a write instruction to modify the data, the cache line is directly operated and the state remains in the modified state.

[0040] Step 7: If the executable file has not been executed, execute step 4; otherwise, end this process. Example

[0041] In this embodiment, an instruction conversion memory access conflict optimization method based on instruction stream feature recognition provided by the present invention is adopted, and a low memory access conflict conversion execution of an x86 architecture executable file on an ARM many-core system is realized by modifying a memory access code generation module in a dynamic instruction conversion engine, including the following steps:

[0042] S1. Detect the shared memory area of ​​the target program based on instruction stream feature detection. The specific steps are as follows:

[0043] S1.1. Identify memory allocation related function calls during dynamic instruction conversion. Under the x86 architecture, for the underlying system calls related to shared memory creation, the assembly level calls are as follows:

[0044] moveax, <shmgetsyscallnumber> / / Put the system call number into the eax register

[0045] move bx, <key> / / Put the shared memory key value into the ebx register to identify the shared memory

[0046] movecx, <size> / / Shared memory size to be placed

[0047] movedx, <flags> / / Put flag bits, such as relevant settings for permissions, etc.

[0048] int 0x80 / / Trigger a software interrupt and enter kernel mode to execute a system call

[0049] After the ARM architecture conversion, the implementation of the ARM system call corresponding to shmget:

[0050] mov r7, <shmgetsyscallnumberforarm> / / Put the system call number of shmget under ARM into register r7

[0051] mov r0, <key> / / Put the shared memory key value into register r0

[0052] mov r1, <size> / / Store the shared memory size in register r1

[0053] mov r2, <flags> / / Put the flag bit into register r2

[0054] svc 0x00 / / Trigger a system call (use the svc instruction in ARM to replace the int instruction in x86)

[0055] Insert the following code after the converted ARM instruction to save the shared memory address. Here, r0 is the return value after svc 0x00:

[0056] str r0, [shared_memory_id] / / Subsequent code can read the return value from shared_memory_id at any time for analysis

[0057] Establish a mapping between the shared memory ID and the shared memory size. Here, the shared memory ID is stored at the address shared_memory_id, and the shared memory size is stored in register r1.

[0058] S1.2. Track memory mapping-related operations. Under the x86 architecture, after obtaining the shared memory segment identifier, it needs to be mapped to the process's address space, which is usually achieved by calling the shmat function. Its assembly-level call is as follows:

[0059] move eax, <shmatsyscallnumber> / / Put the system call number into eax

[0060] move bx, <shmid> / / Put the shared memory segment identifier obtained previously into ebx

[0061] movecx, <addr> / / Write the address of the desired mapping to ecx

[0062] move dx, <flags> / / Put the mapping-related flag bits into edx

[0063] int 0x80 / / Execute a system call for memory mapping

[0064] The instruction after ARM architecture conversion is:

[0065] mov r7, <shmatsyscallnumberforarm> / / Put the shmat system call number under ARM into r7

[0066] mov r0, <shmid> / / Place the shared memory segment identifier into r0

[0067] mov r1, <addr> / / Load the expected mapping address into r1

[0068] mov r2, <flags> / / Put the mapping flag bit into r2

[0069] svc0x00 / / Trigger the system call to complete the mapping

[0070] Insert the following code after the converted ARM instruction to save the shared memory address. Here, r0 is the return value after svc0x00:

[0071] str r0, [shared_memory_address] / / Subsequent code reads the return value from shared_memory_address for analysis

[0072] Based on the mapping of the shared memory ID and shared memory size established in the previous step, establish the mapping of the shared memory address and shared memory size.

[0073] S1.3. Construct a shared memory list denoted as sharedMemoryList through the established mapping. Each element of the list includes the starting address and length of a shared memory, as well as the ID of the process that created this shared memory area.

[0074] Existing shared memory detection methods mainly rely on identifying specific system call function names to determine the creation operation of shared memory. However, in the scenario of dynamic instruction conversion, function names may undergo complex conversions, obfuscations, or optimizations, resulting in the failure of traditional methods. The present invention proposes an intelligent recognition method based on instruction flow characteristics. By analyzing specific patterns and register operation characteristics in the instruction flow and identifying at the level of the lowest-level system calls, it is possible to more accurately determine whether there is a shared memory creation operation.

[0075] S2. Track the shared memories in sharedMemoryList to determine the memories that will actually be accessed by other processes. In S1, it can only be judged whether the memory is used as shared memory but cannot determine whether it will be accessed by other processes. However, only those memory spaces that will be accessed by other processes may have memory access conflicts. Therefore, further judgment and identification are still needed.

[0076] The specific steps are as follows:

[0077] S2.1. Traverse all processes in sharedMemoryList. For each process, denote the ID of the process as processID, obtain the page table entry of its shared memory page, and copy this page table entry and denote it as originalPageTable.

[0078] S2.2. For the shared memory pages recorded in sharedMemoryList, release their mapping relationship to the physical memory pages, thereby releasing the mapping of all processes to them, and then use the page fault exception to track and record the access process, release the page mapping, and hold mmap_lock, as follows:

[0079] pte_t*pte=get_pte(target_vma->vm_mm,target_addr)

[0080] pte_clear(target_vma->vm_mm,target_addr,pte)

[0081] flush_tlb_page(target_vma,target_addr)

[0082] S2.3. Traverse all processes in sharedMemoryList, and for each process with ID processID, copy originalPageTable back to the page table of process processID, that is, restore the process page table that was released in the previous step, so that the page fault exception will be triggered only when other processes access the shared memory in sharedMemoryList.

[0083] S2.4. When a page fault exception is triggered when the process accesses the page, if the address where the exception occurs is in sharedMemoryList, this memory page is counted into usedSharedMemoryList; the mapping is restored to handle the current exception and enable the current process to continue execution normally. The specific examples are as follows:

[0084] set_pte_at(vmf->vma->vm_mm,vmf->address,vmf->pte,orig_pte) / / Reset the original PTE

[0085] structpage*page=pte_page(orig_pte) / / Increase the reference count of the physical page (to prevent it from being released)

[0086] get_page(page) / / Increase refcount

[0087] if(vmf->flags&FAULT_FLAG_WRITE) / / Mark page as accessed / dirty

[0088] pte = pte_mkdirty (orig_pte) / / refresh TLB, optional, set_pte_at may be implicitly processed

[0089] flush_tlb_page(vmf->vma, vmf->address)

[0090] return VM_FAULT_NOPAGE / / Inform the kernel that the mapping has been restored

[0091] S3. For the shared memory that is determined to be accessed by multiple processes, i.e., the shared memory recorded in the usedSharedMemoryList, use a conflict avoidance method based on the lock mechanism to ensure exclusive access to the shared memory. The specific steps are as follows:

[0092] S3.1. Implementation and initialization of the lock. In the generated ARM memory access instructions, use atomic operations and a specific memory area as the lock variable. First, allocate a byte in memory as the lock flag and initialize it to 0, indicating the unlocked state. Assume the address of the lock is stored in r0. The following is the ARM assembly code for initializing the lock:

[0093] LDR r1, =0

[0094] STR B r1, [r0]

[0095] S3.2. Detect whether the memory access address is within the range of a shared memory in the established shared memory list. If so, execute the following steps:

[0096] Lock operation. When the ARM computing core needs to access the shared memory area, it will try to acquire the lock. Use the atomic operations LDREXB (Load Exclusive Byte) and STREXB (Store Exclusive Byte) to implement the lock. If the value loaded by LDREXB is 0 and STREXB successfully stores 1, it means that the core has successfully acquired the lock; otherwise, it will retry. The following is the ARM assembly code for the lock operation:

[0097] lock_loop:

[0098] LDREX B r2, [r0]

[0099] CMP r2, #0

[0100] BNE lock_loop

[0101] MOV r3, #1

[0102] STREX B r4, r3, [r0]

[0103] CMP r4, #0

[0104] BNE lock_loop

[0105] LDR r5, [r6] / / Assume the address of the shared data is stored in r6

[0106] ADD r5, r5, #1

[0107] STR r5, [r6]

[0108] Unlock operation. After the operation is completed, the unlock operation is performed by storing 0 to the lock variable. Assume the address of the lock is stored in r0. The following is the ARM assembly code for the unlock operation:

[0109] LDR r1, =0

[0110] STRB r1, [r0]

[0111] LDR r8, =1

[0112] STRB r8, [r7] / / Assume using the memory flag as a semaphore, stored in r7

[0113] The existing lock mechanisms may be based on library functions of high-level languages, while the present invention uses atomic operations of ARM assembly to implement locks, avoiding the additional overhead of high-level languages and enabling more precise control of the lock granularity. It can implement concurrent control of memory access at a lower level, is applicable to fine-grained access control of shared memory in the ARM multi-core environment, and improves the synchronization and security of memory access. At the same time, combining the unlock operation with the subsequent notification mechanism enables more flexible coordination of the operations of multiple cores after unlocking, avoiding the blind waiting of other cores, thereby improving the overall system parallelism.

[0114] S4. Manage the cache state using the cache coherence protocol to avoid inconsistent memory access. The cache stores a copy of the shared memory, and each computing core maintains its own private cache, where there may be copies of the same shared memory. The specific steps are as follows:

[0115] S4.1. Use the cache coherence protocol of ARM, such as MESI, to reduce memory access conflicts. ARM multi-core systems usually follow the MESI protocol to maintain cache coherence. In the conversion engine, ensure that the instructions after program conversion make full use of the advantages of the MESI protocol. For example, when performing memory access, decide the operation according to the MESI state (Modified, Exclusive, Shared, Invalid).

[0116] S4.2. When a core reads data, if the cache state is Invalid, initiate a read request to obtain the data from memory or the cache of other cores, and update the cache state to Shared or Exclusive. Assume the data address is stored in r0. The following is an example of the ARM assembly code during the read operation:

[0117] LDR r1, [r0]

[0118] MRC p15, 0, r2, c0, c0, 0 / / Check cache status

[0119] CMP r2, #0 / / If the status is Invalid, initiate a read request

[0120] BEQ read_request

[0121] B use_data / / If the status is Shared or Exclusive, directly use the data

[0122] read_request: / / Initiate a read request to obtain data from memory or the cache of other cores

[0123] LDR r1, [r0]

[0124] MCR p15, 0, r3, c0, c0, 0 / / Update the cache status to Shared or Exclusive. Here, the cache status is set through co-processor instructions, and the specific instructions vary depending on the hardware

[0125] LDR r5, [r4] / / Assume the monitored address is stored in r4. If the data is modified, it may be necessary to reread

[0126] CMP r5, #1 / / Assume 1 indicates that the data is being modified

[0127] BEQ reread

[0128] B use_data

[0129] reread:

[0130] LDR r1, [r0]

[0131] MCR p15, 0, r6, c0, c0, 0 / / Update the cache status again

[0132] use_data:

[0133] Existing program transformations may not fully consider the importance of cache coherence protocols in memory access conflicts. The present invention utilizes the ARM's MESI protocol at the assembly level to ensure active management of cache states in a multi-core environment, avoiding inconsistent access to the same data by multiple cores, reducing memory access conflicts caused by cache inconsistencies, and improving data access consistency and performance. The consideration of the cache snooping mechanism is added. When reading data, it can not only rely on the current cache state but also check whether other cores are modifying the data, further enhancing the guarantee of data consistency and avoiding data conflicts caused by cache update delays, which is an innovative extension based on the original.

[0134] In summary, the above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.< / flags> < / addr> < / shmid> < / shmatsyscallnumberforarm> < / flags> < / addr> < / shmid> < / shmatsyscallnumber> < / flags> < / size> < / key> < / shmgetsyscallnumberforarm> < / flags> < / size> < / key> < / shmgetsyscallnumber>

Claims

1. A method for optimizing instruction conversion memory access conflicts based on instruction stream feature recognition, characterized in that: The specific steps include: Step 1: pre-execute an executable file in an ARM many-core system in a dynamic instruction conversion manner, use instruction stream analysis to determine a first x86 instruction sequence related to an x86 system call, parse a first ARM instruction sequence converted from the first x86 instruction sequence to determine a shared memory between processes to form a first shared memory list; Step 2, load and execute the executable file in a dynamic instruction conversion mode, if the current instruction to be converted is a memory access instruction and the memory address is in the first shared memory list, convert the current instruction to be converted into an instruction sequence consisting of a lock instruction sequence, a second ARM instruction sequence and an unlock instruction sequence, and execute step 3 when the current instruction to be converted is a read operation, and execute step 4 when it is a write operation; otherwise, execute step 5; Step 3: If the cache line corresponding to the current computing core is in an invalid state, a read request is sent to the bus to read the data and the state of the cache line is set to a shared state or an exclusive state after the read is completed; if it is a shared state or an exclusive state, data is read from the cache line; and then step 5 is executed; Step 4: If the cache line corresponding to the current computing core is in an exclusive state, the data in the cache line is modified and set to a modified state; if it is in a shared state, an invalidation message is sent to other computing cores, and after the response is confirmed, it is set to a modified state, and then a write operation is performed; if it is in an invalid state, a read-exclusive request is sent to the bus, and after obtaining the data, it is set to a modified state, and then a write operation is performed; if it is in a modified state, a write operation is performed; Step 5: If the executable file has not completed execution, execute step 2; otherwise, end this process.

2. The instruction conversion memory access conflict optimization method according to claim 1, characterized in that: The method of using instruction stream analysis to determine the first x86 instruction sequence related to the x86 system call in step 1 is: after using instruction stream analysis to identify the system call number of the shared memory related x86 system call, the first x86 instruction sequence is formed from the instruction containing the system call number to the soft interrupt trigger instruction.

3. The instruction conversion memory access conflict optimization method according to claim 1, characterized in that: The method of determining the shared memory between processes to form the first shared memory list by parsing the first ARM instruction sequence converted from the first x86 instruction sequence in step 1 is: If the first ARM instruction sequence includes a system call for creating or obtaining a shared memory segment, an ARM instruction A for saving a shared memory key value is added after the first ARM instruction sequence; if the first ARM instruction sequence includes a system call for attaching a shared memory segment to an address space of a calling process, an ARM instruction B for saving a shared memory address is added after the first ARM instruction sequence; the process ID, shared memory key value and shared memory address of the current process are obtained, a shared memory list is formed by the shared memory key value, the shared memory address and the process ID for creating the shared memory, and the shared memory list is used as the first shared memory list.

4. The instruction conversion memory access conflict optimization method according to claim 3, characterized in that: Traverse the process ID in the shared memory list as the target process, obtain the page table entries of all shared memory pages of the target process, and record them as original page table entries; release the mapping relationship between the above shared memory page and the physical memory page, and then save the information of the original page table entry into the page table of the target process, and save the shared memory address that triggers the page fault exception and the exception address belongs to the shared memory list into the first shared memory list.

5. The instruction conversion memory access conflict optimization method according to claim 1, characterized in that: The locking instruction sequence in step 2 is composed of an instruction to set a lock variable, an atomic operation LDREXB to load an exclusive byte, and an atomic operation STREXB to store an exclusive byte.

6. The instruction conversion memory access conflict optimization method according to claim 1, characterized in that: The unlocking instruction sequence in step 2 is composed of instructions for modifying lock variables.

7. The instruction conversion memory access conflict optimization method according to claim 1, characterized in that: In the step 2, the second ARM instruction sequence is composed of ARM instructions with the same functions as the current instruction to be converted and other memory access related instructions.

8. The instruction conversion memory access conflict optimization method according to claim 1, characterized in that: The method of sending a read request to the bus to read data in step 3 is to read data from the cache line or memory of other computing cores.

9. The instruction conversion memory access conflict optimization method according to claim 1, characterized in that: The x86 system calls in step 1 include: shmget, shmat, shmdt and shmctl.

Citation Information

Patent Citations

  • ARM many-core-oriented x86 instruction conversion data synchronization optimization method

    CN119576412A

  • ARM many-core instruction conversion memory access optimization method based on reinforcement learning

    CN119690517A