A x86 instruction dynamic translation cache consistency maintenance method for ARM many-core

By using the x86 cache shared tracking module and synchronization thread in the ARM multi-core system, maintaining multi-dimensional shared identity and completing cache updates, the challenge of cache consistency maintenance of x86 programs on the ARM system is solved, and execution reliability and performance are improved.

CN119645421BActive Publication Date: 2025-05-13北京麟卓信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510146994.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-13
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

When dynamically converting x86 programs and running them on ARM multi-core systems, maintaining cache consistency faces challenges, and the existing MESI protocol cannot directly meet the complex needs of cross-architecture conversion scenarios.

Method used

The x86 cache sharing tracking module is used to realize the maintenance of multi-dimensional shared identity, and the cache update operation is completed asynchronously through the x86 cache synchronization thread to ensure the cache consistency of x86 instructions in the ARM multi-core system.

Benefits of technology

It significantly improves the cross-platform execution reliability and execution performance of x86 programs on ARM multi-core systems, ensuring cache consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119645421B_ABST
    Figure CN119645421B_ABST
Patent Text Reader

Abstract

The present invention discloses an x86 instruction dynamic conversion cache consistency maintenance method for ARM many-cores. By constructing a multi-dimensional shared identifier that describes the cache line status of a computing core related to the x86 instruction, when the ARM many-core system loads an executable file through dynamic instruction conversion, the x86 cache shared tracking module sets the multi-dimensional shared identifier of the cache line for the instruction related to the memory access according to the relationship between the x86 instruction and the cache line, and adds the generated update message to the x86 cache update sequence. The x86 cache synchronization thread asynchronously completes the execution of the update operation in the x86 cache update sequence according to the status of the ARM many-core system, thereby ensuring the cache consistency of the x86 program when it is executed in the ARM many-core system, and significantly improving the reliability and execution performance of the x86 program when it is executed across platforms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of computer software development, and in particular relates to an ARM multi-core oriented x86 instruction dynamic conversion cache consistency maintenance method. Background Art

[0002] As the demand for heterogeneous computing grows, using ARM many-core systems to execute traditional x86 programs has become an important way to expand computing capabilities. ARM many-core systems usually refer to systems with more than 64 computing cores. However, the x86 and ARM architectures differ significantly in cache systems, memory access modes, and multi-core coordination mechanisms. In the process of dynamically converting x86 programs and running them on ARM many-core systems, maintaining cache consistency faces unprecedented challenges. Traditional cache consistency protocols, such as ARM's native MESI protocol, can handle cache consistency issues between homogeneous ARM multi-cores, but cannot directly meet the complex and changing requirements in cross-architecture conversion scenarios.

[0003] Typically, each computing core in an ARM many-core system has two L1 caches and one L2 cache. The same memory data may have multiple different copies in the caches of multiple computing cores, which can lead to data inconsistency, including data consistency between computing cores and between DMA and cache. The MESI protocol defines four states for cache lines, including Modified (M), Exclusive (E), Shared (S), and Invalid (I).

[0004] However, when x86 programs are executed in an ARM many-core system by means of dynamic instruction conversion, the complexity of the x86 program may introduce multiple different states. As a result, the existing cache consistency mechanisms such as the MESI protocol cannot meet actual needs, bringing risks to the cross-platform execution of x86 programs. Summary of the invention

[0005] In view of this, the present invention provides an x86 instruction dynamic conversion cache consistency maintenance method for ARM many-cores, which adopts an x86 cache sharing tracking module to realize the maintenance of multi-dimensional shared identifiers, and the x86 cache synchronization thread completes the execution of update operations in the x86 cache update sequence, thereby realizing the cache consistency of x86 instructions when they are executed in an ARM many-core system.

[0006] The present invention provides an ARM multi-core oriented x86 instruction dynamic conversion cache consistency maintenance method, which specifically includes the following steps:

[0007] Step 1: When the ARM many-core system is initialized, storage space is allocated for the multi-dimensional shared identifier in the cache of each ARM computing core, the x86 cache shared tracking module is started, the x86 cache update sequence is created, and the x86 cache synchronization thread is started; the executable file is loaded through dynamic instruction conversion to obtain the current instruction to be converted;

[0008] Step 2: Convert the current instruction to be converted into an ARM instruction with the same function, and record it as the first ARM instruction; if the first ARM instruction is related to memory access, execute step 3, otherwise execute step 6;

[0009] Step 3: The x86 cache sharing tracking module in the current computing core parses the first ARM instruction, obtains the cache line of the current computing core and its sharing status, and sets the multi-dimensional sharing identifier of the cache line in the current computing core;

[0010] Step 4, the x86 cache sharing tracking module sets the state of the cache line according to the type of the current instruction, determines whether to send a cache line modification notification, and each computing core updates the same data block in its local cache line according to the cache line modification notification; when the computing core performs an update operation on the cache line with the set state, the x86 cache sharing tracking module adds the update request to the x86 cache update sequence;

[0011] Step 5, the x86 cache synchronization thread periodically polls the x86 cache update sequence. When an update request is found, the current load of the ARM many-core system is obtained. If the load is lower than the threshold, the update operation is performed and the executed update request is deleted from the x86 cache update sequence; otherwise, the x86 cache update sequence is kept unchanged, and step 5 is performed after waiting for the set time.

[0012] Step 6: If the executable file has completed execution, then the process ends; otherwise, the next x86 instruction in the executable file is selected as the current instruction to be converted, and step 2 is executed.

[0013] Furthermore, in the step 2, the determination method of whether the first ARM instruction is related to the memory access is: determining by obtaining the operation code of the x86 instruction.

[0014] Furthermore, in step 2, the determination method of whether the first ARM instruction is related to memory access is: determining by analyzing the operand of the first ARM instruction.

[0015] Furthermore, in step 4, when the current instruction is a flow control instruction, the method for predicting the cache line to be used is: obtaining the flow control instruction associated code to predict the instruction execution flow, and predicting the cache line to be used according to the instruction execution flow.

[0016] Furthermore, the x86 instruction type identifier occupies 3 bits for storage, the x86 register mask identifier occupies 8 bits for storage, and the x86 shared core mask identifier occupies 64 bits for storage.

[0017] Furthermore, the multi-dimensional sharing identifier includes an x86 instruction type identifier, an x86 register mask identifier and an x86 shared core mask identifier.

[0018] Furthermore, the step 3 of setting the multi-dimensional sharing identifier of the cache line in the current computing core according to the instruction type, operand and cache line sharing status includes: when the operand contains a register and the register will affect other computing cores, setting the x86 register mask identifier of the current computing core to the mask information of the register; determining that the computing core that currently shares the cache line is recorded as a shared core, and setting the flag bit corresponding to the shared core in the x86 shared core mask identifier of the current computing core.

[0019] Furthermore, in step 3, the first ARM instruction is parsed to obtain the instruction type and operand.

[0020] Furthermore, the x86 cache sharing tracking module in step 4 sets the state of the cache line according to the type of the current instruction, determines whether to send a cache line modification notification, and each computing core updates the same data block in its local cache line according to the cache line modification notification, specifically in the following manner:

[0021] The x86 cache sharing tracking module obtains the state of the cache line, and when the current instruction is a read instruction and the cache line is in an exclusive state, sets the cache line to a shared state; when the current instruction is a storage instruction and the cache line is in an exclusive state, sets the cache line to a modified state, writes the modified data in the cache line back to the main memory, and sends a cache line modification notification to other computing cores; when the current instruction is a flow control instruction, predicts the cache line to be used, and sets the cache line to a shared state; when the current instruction is a data load instruction and other computing cores share the cache line, sets the cache line to a shared state; when the current instruction is an atomic operation instruction, uses an instruction sequence consisting of an LDREX instruction, a first ARM instruction, and a STREX instruction to update the first ARM instruction;

[0022] When a computing core receives a cache line modification notification sent by other computing cores, it marks the same data block in its local cache line as invalid and re-reads it from the main memory when the data is needed; before the computing core executes the LDREX instruction, the x86 cache sharing tracking module sets the cache line to an exclusive state and sends an exclusive notification to other computing cores; after the computing core completes the execution of the STREX instruction, if the execution is successful, the x86 cache sharing tracking module sets the cache line to a shared state and sends a sharing notification to other computing cores, otherwise it continues to maintain the exclusive state.

[0023] Furthermore, the cache behavior in the set state in step 4 is a cache line in a shared state and whose x86 shared core mask identifier is not 0. Beneficial Effects

[0024] The present invention constructs a multi-dimensional shared identifier that describes the status of the computing core cache line related to the x86 instruction. When the ARM many-core system loads the executable file through dynamic instruction conversion, the x86 cache shared tracking module sets the multi-dimensional shared identifier of the cache line for the instruction related to the memory access according to the relationship between the x86 instruction and the cache line, and adds the generated update message to the x86 cache update sequence. The x86 cache synchronization thread asynchronously completes the execution of the update operation in the x86 cache update sequence according to the status of the ARM many-core system, thereby ensuring the cache consistency of the x86 program when it is executed in the ARM many-core system, and significantly improving the reliability and execution performance of the x86 program when it is executed across platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A schematic flow chart of an ARM multi-core oriented x86 instruction dynamic conversion cache consistency maintenance method provided by the present invention. DETAILED DESCRIPTION

[0026] The present invention is described in detail below with reference to the accompanying drawings and with reference to the embodiments.

[0027] In an ARM many-core system, each core has its own corresponding cache. Therefore, when multiple cores access and modify shared memory data, cache consistency needs to be maintained. When an ARM program converted from an x86 program is run in an ARM many-core system, it is necessary to ensure that the read and write operations of memory data by different cores meet the program's expectations. For example, in an x86 program, there may be read and write operations on shared variables. In an ARM multi-core environment, a cache consistency protocol (such as the MESI protocol) is required to ensure the correctness of the data processing process.

[0028] The present invention provides an x86 instruction dynamic conversion cache consistency maintenance method for ARM many-cores, the core idea of ​​which is: constructing a multi-dimensional shared identifier that describes the status of a computing core cache line related to the x86 instruction, and when the ARM many-core system loads an executable file through dynamic instruction conversion, the x86 cache sharing tracking module sets the multi-dimensional shared identifier of the cache line for the instruction related to the memory access according to the relationship between the x86 instruction and the cache line, and adds the generated update message to the x86 cache update sequence, and the x86 cache synchronization thread asynchronously completes the execution of the update operation in the x86 cache update sequence according to the status of the ARM many-core system, thereby ensuring the cache consistency when the x86 program is executed in the ARM many-core system.

[0029] The present invention provides an ARM multi-core oriented x86 instruction dynamic conversion cache consistency maintenance method, the specific process is as follows Figure 1 As shown, the specific steps include:

[0030] Step 1: When the ARM many-core system is initialized, storage space is allocated for the multi-dimensional shared identifier in the cache of each ARM computing core, the x86 cache shared tracking module is started, the x86 cache update sequence is created, and the x86 cache synchronization thread is started; the executable file is loaded through dynamic instruction conversion in the ARM many-core system to obtain the current instruction to be converted.

[0031] Among them, the multi-dimensional shared identifier is used to describe the state of the cache line of the computing core, that is, it describes the impact of the x86 instruction on the cache line state, including the x86 instruction type identifier, the x86 register mask identifier and the x86 shared core mask identifier. The x86 cache sharing tracking module is used for real-time maintenance of multi-dimensional shared identifiers in the dynamic conversion of x86 instructions. The x86 cache update sequence is used to record update requests related to data in the shared cache line. The x86 cache synchronization thread is used to asynchronously process update requests in the x86 cache update sequence.

[0032] Specifically, the x86 instruction type identifier is used to identify the type of x86 instruction that causes the cache line to be accessed. The x86 register mask identifier is the mask information of a register that is an operand of the x86 instruction and may affect data access by other computing cores. The x86 shared core mask identifier is used to identify the computing cores that share data in the cache line, with each computing core corresponding to 1 bit.

[0033] In order to further optimize the storage space in the computing core, the x86 instruction type identifier occupies 3 bits of storage, the x86 register mask identifier occupies 8 bits of storage, and the x86 shared core mask identifier occupies 64 bits of storage.

[0034] Step 2: Convert the current instruction to be converted into an ARM instruction with the same function, and record it as the first ARM instruction; if the first ARM instruction is related to memory access, execute step 3; otherwise, execute step 6.

[0035] Among them, the first ARM instruction is related to memory access, which can be determined by obtaining the operation code of the x86 instruction, and can also be determined by analyzing the operand of the first ARM instruction. For example, when its operand is a memory address, it can be determined that the instruction is related to memory access.

[0036] Step 3: The x86 cache sharing tracking module in the current computing core parses the first ARM instruction to obtain the instruction type and operand, and obtains the cache line of the current computing core and its sharing status, and sets the multi-dimensional sharing identifier of the cache line in the current computing core according to the instruction type, operand and cache line sharing status.

[0037] Specifically, the x86 instruction type identifier of the multi-dimensional shared identifier of the current computing core is set according to the instruction type of the x86 instruction, for example, the data loading instruction type identifier is 000, the data storage instruction type identifier is 001, and the arithmetic operation instruction type identifier related to memory access is 010, etc.

[0038] According to the operand of the obtained x86 instruction, when the operand contains a register and the register may affect other computing cores, the x86 register mask identifier of the multi-dimensional sharing identifier of the current computing core is set to the mask information of the register.

[0039] According to the cache line sharing state, the computing core currently sharing the cache line is determined to be recorded as a shared core, and a flag bit corresponding to the shared core is set in the x86 shared core mask identifier of the multi-dimensional sharing identifier of the current computing core.

[0040] Step 4. The x86 cache sharing tracking module obtains the state of the cache line. When the x86 instruction type identifier of the multi-dimensional sharing identifier shows that the current instruction is a read instruction and the cache line is in an exclusive state, the x86 cache sharing tracking module sets the cache line to a shared state; when the x86 instruction type identifier shows that the current instruction is a storage instruction and the cache line is in an exclusive state, the x86 cache sharing tracking module sets the cache line to a modified state, writes the modified data in the cache line back to the main memory, and sends a cache line modification notification to other computing cores; when the x86 instruction type identifier shows that the current instruction is a flow control instruction, obtain the flow The control instruction association code predicts the instruction execution flow, and predicts the cache line that may be used according to the instruction execution flow, and the x86 cache sharing tracking module sets the cache line to a shared state; when the x86 instruction type identifier shows that the current instruction is a data loading instruction and the x86 shared core mask identifier of the multi-dimensional sharing identifier shows that other computing cores share the cache line, the x86 cache sharing tracking module sets the cache line to a shared state; when the x86 instruction type identifier shows that the current instruction is an atomic operation instruction, the first ARM instruction is updated using an instruction sequence consisting of an LDREX instruction, a first ARM instruction, and a STREX instruction;

[0041] At the same time, when the computing core receives the cache line modification notification sent by other computing cores, it marks the same data block in its local cache line as invalid, and re-reads it from the main memory when the data is needed; before the computing core executes the LDREX instruction, the x86 cache sharing tracking module sets the cache line to an exclusive state and sends an exclusive notification to other computing cores; after the computing core completes the execution of the STREX instruction, if the execution is successful, the x86 cache sharing tracking module sets the cache line to a shared state and sends a sharing notification to other computing cores, otherwise it continues to maintain an exclusive state;

[0042] When the computing core performs an update operation on a cache line that is in a shared state and whose x86 shared core mask flag is not 0, the x86 cache sharing tracking module adds the update request to the x86 cache update sequence.

[0043] Step 5. The x86 cache synchronization thread periodically polls the x86 cache update sequence. When an update request is found, the current load of the ARM many-core system is obtained. If the load is lower than the threshold, the update operation is performed and the executed update request is deleted from the x86 cache update sequence; otherwise, the x86 cache update sequence is kept unchanged, and step 5 is executed after waiting for the set time.

[0044] Step 6: If the executable file has completed execution, then this process ends; otherwise, the next x86 instruction in the executable file is selected as the current instruction to be converted, and step 2 is executed. Example

[0045] In this embodiment, a cache consistency maintenance method for dynamic conversion of x86 instructions for ARM many-core systems provided by the present invention is adopted to implement the execution of x86 programs in the ARM many-core system with guaranteed cache consistency, including the following steps:

[0046] x86 Cache Share Tracking Module - "X86_CACHE_SHARE_TRACKER"

[0047] x86 cache update queue - "X86_CACHE_UPDATE_QUEUE"

[0048] x86 cache synchronization thread - "X86_CACHE_SYNC_THREAD"

[0049] S1. Expand the fine-grained shared identification for the ARM instruction conversion engine and establish a multi-dimensional shared identification to achieve dynamic tracking.

[0050] By extending the cache line management engine of the ARM computing core in the dynamic instruction conversion engine, a new multi-dimensional shared flag is implemented. In addition to the conventional bits indicating the cache line status, the following flags for x86 program conversion scenarios are added:

[0051] The X86_INST_TYPE bit field is used as the x86 instruction type identifier: it occupies 3 bits and is used to identify the type of x86 instruction that causes the cache line to be accessed. For example, 000 indicates a data loading instruction, such as MOV reg, [mem]; 001 indicates a data storage instruction, such as MOV [mem], reg; 010 indicates an instruction that performs arithmetic operations and the operand involves memory access, such as ADDreg, [mem]. By accurately recording the instruction type, cache consistency issues caused by different types of instructions can be handled in a targeted manner.

[0052] Use the X86_SRC_REG_MASK bit segment as the x86 register mask identifier: When the x86 instruction operand involves a register and the register content may affect the data access of other computing cores, use this bit segment to record the register mask information. For example, in the instruction sequence MOV EAX, [mem]ADD EBX, EAX, if the instruction sequence causes a cache line to be accessed, X86_SRC_REG_MASK will record the mask corresponding to the EAX register, so as to judge the data flow and potential sharing risks.

[0053] The X86_SHARED_CORE_MASK bit field is used as the x86 shared core mask identifier: 64 bits correspond to a maximum of 64 possible ARM computing cores, and can also be set to other values ​​to clearly identify which ARM computing cores share the data of the current cache line. Each bit corresponds to an ARM computing core, and the initial value is all 0. When a computing core accesses the cache line due to executing the converted x86 instruction, its corresponding bit is set to 1. This embodiment reads the information of the computing cores that share the cache line through the MESI protocol.

[0054] Here is an example of x86 assembly code and the multi-dimensional shared flags it sets when converted to ARM instructions:

[0055] The instructions in an x86 program are:

[0056] MOV EAX, [memory address 0x2000]

[0057] ADD EAX, EBX

[0058] MOV [memory address 0x2000], EAX

[0059] The converted ARM instructions are:

[0060] LDR R0, =0x2000

[0061] LDR R1, [R0] / / Read memory data to R1, corresponding to the x86 read operation

[0062] At this time, the X86_INST_TYPE of the cache line is set to 000, indicating a data loading instruction; X86_SRC_REG_MASK records the mask corresponding to R1 based on the mapping of the EAX register in the ARM system, assuming that the register is mapped to R1; X86_SHARED_CORE_MASK sets the bit corresponding to the currently executing computing core to 1.

[0063] ADD R1, R2 / / Assume EBX is mapped to R2 and perform addition

[0064] STR R1, [R0] / / Write back data, corresponding to x86 write operation

[0065] Update the value of X86_INST_TYPE to 001, indicating a data storage instruction; keep the value of X86_SRC_REG_MASK unchanged. Since the write operation may affect the shared data of other computing cores, broadcast the cache line update information to the computing cores, which can be achieved through the system bus or a specific inter-core communication channel, and update X86_SHARED_CORE_MASK at the same time. If other computing cores respond and confirm that the corresponding shared cache line exists in their cache, the corresponding bit of the information sender of their own X86_SHARED_CORE_MASK is set to 1.

[0066] S2. Establish X86_CACHE_SHARE_TRACKER as the x86 cache sharing tracking module, running in each ARM computing core, responsible for real-time tracking and updating the above multi-dimensional sharing identifier.

[0067] When executing a converted x86 instruction involving memory access, the instruction first enters X86_CACHE_SHARE_TRACKER for preprocessing. The module dynamically updates the multi-dimensional sharing identifier of the cache line based on the instruction operands, register usage, and current program execution flow information. For example, if an x86 instruction is executed multiple times in a loop body, X86_CACHE_SHARE_TRACKER will re-evaluate the data dependency each time it is executed. If new potential sharing situations are found, such as the introduction of new registers to participate in memory operations during loop iterations, the multi-dimensional sharing identifier will be adjusted in a timely manner. The specific steps are as follows:

[0068] S2.1. Analyze instruction operands.

[0069] For each assembly instruction to be executed, first parse its operands. For example, in ARM assembly, for the instruction LDR R0, [R1], R1 is the source operand and R0 is the destination operand. Determine whether the operand is stored in a register or in memory. In the above example, the source operand is a memory address indirectly addressed by register R1. If the operand involves a memory address, then find the cache line corresponding to the memory address. Among them, the method of finding the cache line corresponding to the memory address can be: find the corresponding cache line through the cache mapping mechanism, such as direct mapping, set associative mapping, etc.

[0070] S2.2. Update the shared identifier of the cache line.

[0071] According to the above information, the shared identifier of the cache line involved is updated.

[0072] If the instruction is a read operation, such as LDR R0, [R1], and the corresponding cache line is in an exclusive state, it is updated to a shared state, and multiple computing cores may read the data.

[0073] For storage operations, such as STR R0, [R1], if the operand is exclusive, it may be necessary to mark it as modified and notify other cores that the cache line has been modified through a cache coherence protocol, such as the MESI protocol.

[0074] For branch and jump instructions, the shared flags of the cache lines that may be involved are updated in advance according to the predicted execution flow. For example, if it is predicted that a function call will access a global variable, the cache line corresponding to the global variable is marked as potentially shared.

[0075] S2.3. Implementation of the consistency protocol.

[0076] When updating the shared flag of a cache line, a cache consistency protocol, such as the MESI protocol, must be followed.

[0077] If the cache line is marked as modified, it needs to be written back to the main memory and the other computing cores are notified that the cache line is invalid. For cache lines in the shared state, when modification notifications are received from other computing cores, they need to be marked as invalid and the latest data needs to be retrieved from the main memory or the cache of other computing cores when necessary.

[0078] The following is an ARM assembly code example:

[0079] MOV R1, #addr / / Store the address value in R1

[0080] LDR R0, [R1] / / Read data from the address pointed to by R1 to R0

[0081] STR R0, [R2] / / Store the value of R0 to the address pointed to by R2

[0082] For the MOV R1, #addr instruction, the operand is parsed first, that is, the immediate value #addr is stored in R1. No memory operation is involved at this time, and the shared flag of the cache line is not affected for the time being.

[0083] For the LDR R0, [R1] instruction, the corresponding memory address is found according to R1, assuming that the address is in the cache. If the cache line was previously in an exclusive state, it will be updated to a shared state due to the read operation, because other computing cores may also read the data.

[0084] For the STR R0, [R2] instruction, the value of R0 is stored at the address pointed to by R2. If the cache line was previously in a shared state, it needs to be updated to a modified state and the other computing cores are notified that the cache line has been modified. A consistency protocol is needed to ensure that the cache lines of other computing cores are invalidated or updated.

[0085] The present invention dynamically adjusts the multi-dimensional shared identifier of the cache line according to the execution of the assembly instruction to maintain the cache consistency in the ARM many-core system. Through the comprehensive analysis of the instruction operands, register usage and program execution flow, the cache consistency problem can be predicted and processed more accurately, and the system performance and data consistency can be improved. This dynamic update mechanism can reduce unnecessary cache failures and data transmission, and optimize the overall performance of the system.

[0086] When the dynamic instruction conversion engine receives an x86 program, it strictly follows the above steps to set the shared flag when converting x86 instructions to ARM instructions. At the same time, the conversion engine works closely with the X86_CACHE_SHARE_TRACKER module to ensure that the shared state of the cache line is accurately identified before the instruction enters the execution phase. By dynamically tracking the shared state at runtime, it adapts to the complex and changeable execution characteristics of x86 programs and overcomes the limitations of traditional static identification when facing dynamic program behavior.

[0087] S3. Implement an automated cache state transition and synchronization mechanism based on x86 instruction-aware cache state migration, that is, expand the ARM native cache state machine in the dynamic execution conversion engine to make it compatible with x86 program conversion requirements. In addition to following the conventional MESI protocol state transition rules, a new specific conversion path for x86 instructions is added:

[0088] When executing an x86 data load instruction and the X86_SHARED_CORE_MASK of the corresponding cache line shows that other computing cores share the data, even if the local cache line is currently in the Exclusive state, it must be forced to be converted to the Shared state to prevent subsequent write operations from destroying the data consistency of other computing cores. For example, if the local cache line was originally in the Exclusive state, LDR R1, [R0] (corresponding to the x86 read operation) is executed and it is detected that other computing cores share it, a signal is immediately sent to the memory controller to update the cache line state to Shared, and the local cache line state record is updated.

[0089] For x86 atomic operation instructions, such as instructions with the LOCK prefix, special cache state control sequences are introduced when converting to ARM instructions. For example, for the x86 atomic exchange and addition instruction LOCK XADD [lock address], ECX, it is converted to the following ARM instruction:

[0090] LDREX R1, [lock address]

[0091] ADD R1, R0

[0092] STREX R2, R1, [lock address]

[0093] At the same time, before executing the LDREX instruction, the local cache line state is forced to be converted to the Exclusive state, and a signal is sent to other computing cores to prevent them from accessing the memory area corresponding to the cache line during the atomic operation. After the atomic operation is completed, according to the return result of the STREX instruction, success or failure, it is decided whether to restore the cache line state to Shared or keep the Exclusive state, and the relevant computing cores are notified to update their cache states.

[0094] S3.2. Implement an asynchronous consistency synchronization strategy, that is, implement an asynchronous cache consistency synchronization mechanism to reduce the impact of synchronization operations on program execution performance. The specific steps are as follows:

[0095] When the ARM computing core needs to update the shared cache line data, it can be indicated by X86_SHARED_CORE_MASK and X86_INST_TYPE that it will no longer immediately perform synchronization operations such as writing back to the memory and notifying other computing cores to update, but instead put the update request into the x86 cache update sequence X86_CACHE_UPDATE_QUEUE.

[0096] Each ARM computing core runs the x86 cache synchronization thread X86_CACHE_SYNC_THREAD in the background and polls the queue regularly. When an update request is found, the thread selects the optimal synchronization time and strategy based on the current cache line status, sharing flag and system load. The specific steps are as follows:

[0097] In the initialization phase of the ARM many-core system, a storage space with a newly added multi-dimensional shared identifier is configured for the cache of each ARM computing core, the x86 cache sharing tracking module X86_CACHE_SHARE_TRACKER is started, the X86_CACHE_UPDATE_QUEUE is created, and the X86_CACHE_SYNC_THREAD is started;

[0098] If the system load is light, for example, the CPU utilization is less than 50%, and most of the relevant computing cores are idle, the synchronization operation is performed immediately, that is, writing back to the memory and sending update notifications to other computing cores through inter-core communication, while updating their cache line status.

[0099] If the system is under high load, delay synchronization and keep the update request in the queue until the right time. For example, when the end of a program phase is detected, such as a function call return, a large loop exit, or a system idle period, batch process the update requests in the queue to improve synchronization efficiency.

[0100] During program execution, the X86_CACHE_SHARE_TRACKER of each computing core continuously monitors instruction execution and updates the shared identifier in real time; X86_CACHE_SYNC_THREAD asynchronously processes cache update requests based on system load and program stage, maintains cache consistency of the entire ARM many-core system, and ensures stable and efficient operation of the converted x86 program.

[0101] The present invention adopts a combination of asynchronous processing and scheduling to flexibly cope with complex operating scenarios of the ARM many-core system and maximize the overall system performance while ensuring cache consistency.

[0102] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for maintaining x86 instruction dynamic conversion cache consistency for ARM multi-core, characterized in that: The specific steps include: Step 1: When the ARM many-core system is initialized, storage space is allocated for the multi-dimensional shared identifier in the cache of each ARM computing core, the x86 cache shared tracking module is started, the x86 cache update sequence is created, and the x86 cache synchronization thread is started; the executable file is loaded through dynamic instruction conversion to obtain the current instruction to be converted; Step 2, convert the current instruction to be converted into an ARM instruction with the same function, and record it as the first ARM instruction; determine whether the first ARM instruction is related to the memory access, if so, execute step 3, otherwise execute step 6; Step 3: In the current computing core, the x86 cache sharing tracking module parses the first ARM instruction, obtains the cache line of the current computing core and its sharing status, and sets the multi-dimensional sharing identifier of the cache line in the current computing core; Step 4, the x86 cache sharing tracking module sets the state of the cache line according to the type of the current instruction, determines whether to send a cache line modification notification, and each computing core updates the same data block in its local cache line according to the cache line modification notification; when the computing core performs an update operation on the cache line with the set state, the x86 cache sharing tracking module adds the update request to the x86 cache update sequence; Step 5, the x86 cache synchronization thread periodically polls the x86 cache update sequence. When an update request is found, the current load of the ARM many-core system is obtained. If the load is lower than the threshold, the update operation is performed and the executed update request is deleted from the x86 cache update sequence; otherwise, the x86 cache update sequence is kept unchanged and step 5 is re-executed after waiting for a set time. Step 6: If the executable file has completed execution, then this process ends; otherwise, the next x86 instruction in the executable file is selected as the current instruction to be converted, and step 2 is executed.

2. The x86 instruction dynamic conversion cache consistency maintenance method according to claim 1, characterized in that: The method for determining whether the first ARM instruction is related to memory access in step 2 is: determining by obtaining the operation code of the x86 instruction.

3. The x86 instruction dynamic conversion cache consistency maintenance method according to claim 1, characterized in that: In the step 2, the determination method of whether the first ARM instruction is related to the memory access is: determining by analyzing the operand of the first ARM instruction.

4. The x86 instruction dynamic conversion cache consistency maintenance method according to claim 1, characterized in that: The multi-dimensional sharing identifier includes an x86 instruction type identifier, an x86 register mask identifier and an x86 shared core mask identifier.

5. The x86 instruction dynamic conversion cache consistency maintenance method according to claim 4, characterized in that: The x86 instruction type identifier occupies 3 bits for storage, the x86 register mask identifier occupies 8 bits for storage, and the x86 shared core mask identifier occupies 64 bits for storage.

6. The x86 instruction dynamic conversion cache consistency maintenance method according to claim 5, characterized in that: The step 3 of setting the multi-dimensional sharing identifier of the cache line in the current computing core includes: when the operand contains a register and the register will affect other computing cores, setting the x86 register mask identifier of the current computing core to the mask information of the register; determining that the computing core that currently shares the cache line is recorded as a shared core, and setting the flag bit corresponding to the shared core in the x86 shared core mask identifier of the current computing core.

7. The x86 instruction dynamic conversion cache consistency maintenance method according to claim 1, characterized in that: In step 3, the first ARM instruction is parsed to obtain the instruction type and operand.

8. The x86 instruction dynamic conversion cache consistency maintenance method according to claim 1, characterized in that: In step 4, the x86 cache sharing tracking module sets the state of the cache line according to the type of the current instruction, determines whether to send a cache line modification notification, and each computing core updates the same data block in its local cache line according to the cache line modification notification, specifically in the following manner: The x86 cache sharing tracking module obtains the state of the cache line, and when the current instruction is a read instruction and the cache line is in an exclusive state, sets the cache line to a shared state; when the current instruction is a storage instruction and the cache line is in an exclusive state, sets the cache line to a modified state, writes the modified data in the cache line back to the main memory, and sends a cache line modification notification to other computing cores; when the current instruction is a flow control instruction, predicts the cache line to be used, and sets the cache line to a shared state; when the current instruction is a data load instruction and other computing cores share the cache line, sets the cache line to a shared state; when the current instruction is an atomic operation instruction, uses an instruction sequence consisting of an LDREX instruction, a first ARM instruction, and a STREX instruction to update the first ARM instruction; When a computing core receives a cache line modification notification sent by other computing cores, it marks the same data block in its local cache line as invalid and re-reads it from the main memory when the data is needed; before the computing core executes the LDREX instruction, the x86 cache sharing tracking module sets the cache line to an exclusive state and sends an exclusive notification to other computing cores; after the computing core completes the execution of the STREX instruction, if the execution is successful, the x86 cache sharing tracking module sets the cache line to a shared state and sends a sharing notification to other computing cores, otherwise it continues to maintain the exclusive state.

9. The x86 instruction dynamic conversion cache consistency maintenance method according to claim 8, characterized in that: The method of predicting the cache line to be used when the current instruction is a flow control instruction in step 4 is: obtaining the flow control instruction associated code to predict the instruction execution flow, and predicting the cache line to be used according to the instruction execution flow.

10. The x86 instruction dynamic translation cache consistency maintenance method according to claim 1, characterized in that: The cache behavior in the set state in step 4 is in a shared state and the cache line whose x86 shared core mask identifier is not 0.

Citation Information

Patent Citations

  • Implementation method and device for cache consistency of multi-core processor, the multi-core processor and storage medium

    CN112416615A

  • Method and device for verifying cache consistency of multi-core processor

    CN116167310A