Real-time Simulation Performance Evaluation Method for Multi-core Processors Based on ARMv8 Architecture
By setting up a concurrent monitoring module in the ARMv8 architecture multi-core processor, collecting and parsing access instruction sequences and memory barrier instructions, identifying concurrency consistency problems, the accuracy of instruction rearrangement in multi-core concurrent operations is solved, the reliability and accuracy of real-time simulation performance evaluation is improved, and a highly targeted evaluation tool is provided.
Patent Information
- Application Number
- CN202510006599.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The prior art is difficult to accurately identify and analyze potential consistency problems caused by instruction rearrangement in multi-core concurrent operations in ARMv8 architecture multi-core processors, affecting the accuracy and reliability of real-time simulation performance evaluation.
By setting up a concurrency monitoring module, the core access instruction sequence and memory barrier instruction information are collected and processed, the address cross-relationship and timing dependencies are analyzed, the instruction rearrangement risk access group is generated, and the high-intensity concurrent operations are injected into the actual simulation environment, the read and write operations of shared variables are tracked, the execution sequence matching degree of the instruction rearrangement risk access group is statistically determined, and the real-time simulation performance evaluation indicators are generated.
It improves the accuracy and applicability of real-time simulation performance evaluation of multi-core processors, can identify potential race conditions in high-intensity concurrency scenarios, provides reliable performance evaluation tools, guides memory consistency optimization, and expands the scope of real-time simulation application under weak memory models.
Smart Images

Figure CN119847896B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-core performance evaluation, and more specifically, to a real-time simulation performance evaluation method for multi-core processors based on the ARMv8 architecture. Background Art
[0002] The ARMv8 architecture adopts a weak memory model to achieve high-performance and low-power optimization. This model allows out-of-order execution and memory read / write reordering. This memory model significantly improves the parallel computing ability in multi-core processors, but at the same time increases the complexity of accessing shared variables. Especially in real-time simulation applications, the high-frequency access of multiple cores to shared variables may lead to unpredictable execution sequences, thereby potentially affecting the stability and consistency of simulation results. Although existing technologies have tried to ensure memory consistency by using means such as "locks" or "memory barriers", in the face of complex instruction reordering problems in multi-core high-concurrency environments, the reliability and applicability of existing methods still have certain limitations.
[0003] In the prior art, how to accurately identify and analyze potential consistency problems caused by instruction reordering in multi-core concurrent operations under the weak memory model will significantly affect the accuracy and reliability of real-time simulation performance evaluation. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a real-time simulation performance evaluation method for multi-core processors based on the ARMv8 architecture to solve the problems raised in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A real-time simulation performance evaluation method for multi-core processors based on the ARMv8 architecture, comprising the following steps:
[0007] S1: Set a concurrent monitoring module in the multi-core processor to collect the access instruction sequences of each processing core to shared variables and the corresponding memory barrier instruction information;
[0008] S2: Analyze the collected access instruction sequences and memory barrier instruction information, and generate instruction reordering risk access groups based on the address cross-relationship and timing dependence between access instructions;
[0009] S3: Conduct concurrent consistency analysis on the instruction reordering risk access groups, identify race conditions that may occur in the use of low-level hardware primitives, and record the synchronization methods corresponding to the access instruction sequences and the distribution of memory barrier instructions;
[0010] S4: Inject a high-intensity concurrent operation scenario in the actual simulation environment, use the concurrent monitoring module to continuously track the read and write operations of shared variables, and compare the execution order of each instruction reordering risk access group with the expected order obtained by parsing;
[0011] S5: Statistically analyze the execution sequence matching degree of each instruction reordering risk access group under different simulation load conditions, and comprehensively generate real-time simulation performance evaluation indicators for the weak memory model condition.
[0012] In a preferred embodiment, a concurrent monitoring module is set in the multi-core processor to collect the access instruction sequences of each processing core to shared variables and the corresponding memory barrier instruction information, specifically including:
[0013] Configure the concurrent monitoring module to record the access operations of each core in the multi-core processor to shared variables, collect access instruction information, and the access instruction information includes the operation type, target address, and the time when it occurs;
[0014] In the concurrent monitoring module, capture memory barrier instructions in real time, extract the type of memory barrier instructions, the specific insertion location, and their association relationship with access instructions;
[0015] Summarize the access instruction information and memory barrier instruction information collected by the concurrent monitoring module, and perform time correction on the collected data based on a unified time reference to generate a global record containing the access instruction sequences of all processing cores and memory barrier instruction information.
[0016] In a preferred embodiment, analyze the collected access instruction sequences and memory barrier instruction information, and generate instruction reordering risk access groups based on the address cross-relationship and timing dependence between access instructions, specifically including:
[0017] Analyze the collected access instruction sequences, extract the target address, operation type, and occurrence time of each access instruction, and sort the access instructions of the same core according to the occurrence time to form an in-core access sequence;
[0018] Based on the in-core access sequence obtained by parsing, identify the cross-relationship of target addresses, determine whether multiple cores simultaneously access the memory address of the same shared variable, and record cross-access instruction pairs;
[0019] Analyze the time information of the cross-access instruction pairs, and judge whether there is timing dependence according to the occurrence time of the access instructions and their association with memory barrier instructions, and record instruction pairs with time-dependent relationships;
[0020] Combine cross-access instruction pairs and instruction pairs with temporal dependencies, and generate an instruction reordering risk access group according to the core number, access target address, and time order.
[0021] In a preferred embodiment, perform concurrent consistency analysis on the access instruction reordering risk access group, identify potential race conditions that may occur during the use of low-level hardware primitives, and record the synchronization methods and memory barrier instruction distributions corresponding to the access instruction sequences, specifically including:
[0022] Classify the access instructions in the access instruction reordering risk access group. According to the operation type, target address, and occurrence time of the access instructions, classify the access instructions in the access instruction reordering risk access group into a read access instruction group, a write access instruction group, and a mixed operation access instruction group;
[0023] For the access instructions in each group, analyze the usage of low-level hardware primitives, including marking the dependency relationships, scopes of action, and execution orders of the hardware access instructions;
[0024] Based on the grouping, analyze whether there are potential race conditions for the access instructions on the same target address according to the target address and time relationships between the access instructions, including the situation where multiple cores simultaneously access the same target address and the memory barrier instructions do not cover all access instructions;
[0025] Record the synchronization methods involved in the access instruction sequences of each group, specifically including the type of synchronization method, the synchronization scope, and the specific access instructions on which it acts.
[0026] In a preferred embodiment, inject a high-intensity concurrent operation scenario in the actual simulation environment, use the concurrent monitoring module to continuously track the read and write operations of shared variables, and compare the execution order of each instruction reordering risk access group with the expected order obtained by parsing, specifically including:
[0027] Construct an actual simulation environment, set a multi-core high-concurrency operation scenario, including defining the target address range of shared variables and allocating the access operation tasks of each core to ensure that multiple cores simultaneously execute read and write operations;
[0028] Inject a high-intensity concurrent load into the simulation scenario, and set each core to access the shared variables at randomized time intervals;
[0029] Use the concurrent monitoring module to continuously track all read and write operations of the shared variables, record the operation type, target address, and execution time of each access, and generate a complete access instruction sequence;
[0030] Compare the access instruction sequence of the trace record with the expected order obtained by parsing, and analyze whether there is instruction reordering in the actual execution order;
[0031] Store the detected instruction reordering situations in the comparison result in the form of structured data, and record the cause, scope of influence, and access instruction numbers involved in each instruction reordering.
[0032] In a preferred embodiment, count the execution sequence matching degrees of each instruction reordering risk access group under different simulation load conditions, and comprehensively generate real-time simulation performance evaluation indicators for the weak memory model, specifically including:
[0033] Set multiple load conditions in the simulation environment. The load conditions include high-intensity load, medium-intensity load, and low-intensity load. Under each load condition, different concurrent scenarios are formed by adjusting the access frequency and operation interval of multiple cores to shared variables;
[0034] For each instruction reordering risk access group under each load condition, compare group by group according to the actual execution order and the expected order of the access instruction sequence, and record the execution sequence matching degree of each group of access instructions. The execution sequence matching degree is determined by comparing whether the actual execution order is consistent with the expected order obtained by parsing;
[0035] Based on the matching degree data of multiple load conditions, generate real-time simulation performance evaluation indicators for the weak memory model. The real-time simulation performance evaluation indicators include the execution consistency situation of each instruction reordering risk access group under different load conditions, the frequency of instruction reordering, and the instruction execution time deviation.
[0036] In a preferred embodiment, under high-intensity load conditions, each core frequently accesses shared variables at the minimum time interval to simulate an extremely high-concurrency environment; under medium-intensity load conditions, each core accesses shared variables at a relatively balanced time interval; under low-intensity load conditions, each core accesses shared variables at a larger time interval to simulate a low-concurrency environment;
[0037] Under each load condition, adjust the access time interval of the core to the shared variable by introducing a randomization mechanism; the access order and time interval between cores are generated by a random function.
[0038] The technical effects and advantages of the real-time simulation performance evaluation method based on the ARMv8 architecture multi-core processor of the present invention:
[0039] 1. By setting up a concurrent monitoring module, accurately collect the access instruction sequences and memory barrier instruction information in a multi-core processor, analyze the address cross-relationship and timing dependency between access instructions, and construct an instruction reordering risk access group. This mechanism enables the clear positioning of the dependency relationships and risk areas of critical instructions, laying a solid data foundation for subsequent consistency analysis and performance evaluation. Compared with the prior art, the technical means of the present invention have higher precision and applicability, and can effectively identify potential race conditions especially in high-intensity concurrent operation scenarios.
[0040] 2. Inject high-intensity concurrent operations under various load conditions in an actual simulation environment, dynamically track the instruction execution order and compare it with the expected order obtained by parsing, so as to accurately calculate the matching degree and reordering frequency of instruction execution. Combining the evaluation results of multi-load scenarios, generate real-time simulation performance evaluation indicators for weak memory models, which can comprehensively reflect the execution consistency of each instruction reordering risk access group under different conditions. The method of the present invention provides a highly targeted and data-transparent evaluation tool, which not only improves the reliability of simulation performance evaluation, but also provides a guiding basis for the memory consistency optimization of multi-core processors, significantly expanding the application scope of real-time simulation under weak memory models. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a schematic diagram of the real-time simulation performance evaluation method for a multi-core processor based on the ARMv8 architecture according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0043] Embodiment: Figure 1 A real-time simulation performance evaluation method for a multi-core processor based on the ARMv8 architecture according to the present invention is given, which includes the following steps:
[0044] S1: Set up a concurrent monitoring module in the multi-core processor to collect the access instruction sequences of each processing core to shared variables and the corresponding memory barrier instruction information.
[0045] S2: Analyze the collected access instruction sequences and memory barrier instruction information, and generate an instruction reordering risk access group according to the address cross-relationship and timing dependency between access instructions.
[0046] S3: Perform concurrent consistency analysis on the instruction reordering risk access groups, identify the race conditions that may occur during the use of low-level hardware primitives, and record the synchronization methods and the distribution of memory barrier instructions corresponding to the access instruction sequences.
[0047] S4: Inject a high-intensity concurrent operation scenario in the actual simulation environment, use the concurrent monitoring module to continuously track the read and write operations of shared variables, and compare the execution order of each instruction reordering risk access group with the expected order obtained by parsing.
[0048] S5: Statistically analyze the execution sequence matching degree of each instruction reordering risk access group under different simulation load conditions, and comprehensively generate real-time simulation performance evaluation indicators for the weak memory model condition.
[0049] Set up a concurrent monitoring module in the multi-core processor to collect the access instruction sequences of shared variables by each processing core and the corresponding memory barrier instruction information, specifically including:
[0050] Configure the concurrent monitoring module to record the access operations of shared variables by each core in the multi-core processor, collect access instruction information, and the access instruction information includes the operation type, target address, and the time when it occurs:
[0051] First, deploy a concurrent monitoring module in the multi-core processor. This module records the access instructions of shared variables involved in the execution of each core by directly accessing the hardware performance counters or instruction trace interfaces of the multi-core processor (such as the ETM of ARM or similar mechanisms). The specific recorded content includes: the operation type of the access instruction, the target address, and the time when the access instruction occurs.
[0052] The operation type indicates the function of the access instruction, such as read (Load), write (Store), or update (Update). This information is obtained by decoding the processor instruction stream and is associated with the specific memory address range of the shared variable.
[0053] The target address records the specific memory location of the access, which is used for subsequent analysis of the cross-relationship between instructions. The target address is extracted from the address bus or cache access records of the processor through hardware probes or software instrumentation.
[0054] The time when it occurs is obtained through the global clock synchronization mechanism between the concurrent monitoring module and the multi-core processor, ensuring the temporal accuracy of the recorded data. The time information is stored with nanosecond-level precision for subsequent analysis of the temporal dependency relationship between instructions.
[0055] Capture memory barrier instructions in real time in the concurrent monitoring module, and extract the type of memory barrier instructions, the specific inserted location, and their association relationship with access instructions:
[0056] The concurrent monitoring module needs to capture in real time the memory barrier instructions executed by each core in a multi-core processor. These barrier instructions are used to coordinate the memory access order between cores and thus are an important basis for analyzing the risk of instruction reordering. The process of capturing memory barrier instructions includes the following aspects:
[0057] Type identification of barrier instructions: By analyzing the binary encoding of the instructions, determine the specific type of the barrier instruction, such as a global barrier (DMB), a data synchronization barrier (DSB), or an instruction synchronization barrier (ISB). These types indicate the scope and priority of the barrier.
[0058] Extraction of insertion position: Record the position of the barrier instruction in the access instruction sequence to clarify its order in the execution sequence. The insertion position is obtained by real-time parsing of instruction pipeline information or instruction trace logs.
[0059] Analysis of associated access instructions: The association relationship between the captured barrier instruction and the preceding and following access instructions is the key to generating an instruction dependency model subsequently. Specifically, it is necessary to determine which access instructions to shared variables the barrier instruction acts on and whether there are cases of coverage or incomplete protection.
[0060] The information captured by the barrier instruction provides complete data input for subsequent parsing and analysis through cross-matching with the access instruction information.
[0061] Summarize the access instruction information and memory barrier instruction information collected by the concurrent monitoring module, and perform time correction on the collected data based on a unified time reference to generate a global record containing the access instruction sequences and memory barrier instruction information of all processing cores:
[0062] Associate and store the access instruction information and memory barrier instruction information according to fields such as core number, timestamp, and target address to form a unified data structure. The data is stored in the form of event logs, and each record contains the instruction type, target address, timestamp, and barrier association information.
[0063] Since each core in a multi-core processor may have an independent local clock, it is necessary to perform time correction on the collected data to generate a globally consistent time reference. The time correction process is based on the global clock source provided by the processor or achieved through an external synchronization signal.
[0064] By sorting the time-corrected data, a complete global access instruction sequence and global memory barrier instruction information are formed. The access instruction sequence is arranged in chronological order, and the memory barrier instructions are inserted as markers into the access instruction sequence to clarify their scope and position of action.
[0065] Store the generated global access records in the storage unit of the concurrency monitoring module in the form of structured data. The storage format includes, but is not limited to, JSON, CSV, or binary log files, for efficient reading during subsequent parsing and analysis. During the storage process, ensure the security and consistency of the data to avoid affecting the accuracy of subsequent steps due to storage latency or data loss.
[0066] Parse the collected access instruction sequence and the memory barrier instruction information. Based on the address cross-relationship and timing dependency among the access instructions, generate an instruction reordering risk access group, specifically including:
[0067] Parse the collected access instruction sequence, extract the target address, operation type, and occurrence time of each access instruction, and sort the access instructions of the same core according to the occurrence time to form an in-core access sequence:
[0068] Extract the target address: For each access instruction, parse its target address information from the access record. The target address refers to the memory location operated by the access instruction and is used to determine whether the access instruction involves shared variables. During parsing, ensure that the target addresses of all instructions are consistent with the collected shared variable address range; otherwise, filter out irrelevant instructions.
[0069] Extract the operation type: The operation type defines the function of the access instruction, including read operation (Load), write operation (Store), and update operation (Update). Extract the operation type information by decoding the instruction opcode and mark it in the parsing result for subsequent analysis.
[0070] Extract the occurrence time: Each access instruction is attached with a timestamp indicating the specific time when the instruction is executed by the processing core. The timestamp is generated by synchronizing the concurrency monitoring module with the global time reference to ensure the consistency of time information among different cores. Parse the timestamp and associate it with the instruction to form the complete time information of the instruction.
[0071] Sort the access instructions of the same core according to the timestamp to generate an in-core access sequence. The sorted access sequence will be used as the basic data for subsequent analysis to identify the address cross-relationship and timing dependency among the instructions.
[0072] Based on the parsed in-core access sequence, identify the cross-relationship of the target addresses, determine whether multiple cores access the memory address of the same shared variable simultaneously, and record the cross-access instruction pairs:
[0073] Determine address cross-relationships: Cross-compare the access instructions of each core with those of other cores to identify whether multiple cores are accessing the same target address simultaneously. The identification of cross-relationships is completed through the matching of target addresses, that is, if the target addresses of two instructions are the same, it is considered that there is an address cross-relationship.
[0074] Once the address cross-relationship is identified, record the relevant pairs of access instructions. Each pair of cross-access instructions contains information about the two instructions, including the core number, target address, operation type, and occurrence time.
[0075] If there are cross-relationships between multiple pairs of instructions between two cores, further filter according to the timestamp, and only retain the pairs of instructions with potential impacts, such as the most recently occurred or cross-relationships covering a larger time window.
[0076] Analyze the time information of the pairs of cross-access instructions. Based on the occurrence time of the access instructions and their relevance to memory barrier instructions, determine whether there is a timing dependency, and record the pairs of instructions with a time-dependent relationship:
[0077] Compare the timestamps of the two instructions in the pair of cross-access instructions and calculate their time difference. If the time difference is below a certain threshold, there may be a potential timing conflict. For example: if the time difference is less than the minimum time unit for instruction processing, it may cause access override; if the time difference exceeds a certain range, it can be regarded as an irrelevant instruction.
[0078] Compare the timestamps of the pairs of cross-access instructions with the insertion time of the memory barrier instructions. If the pair of cross-access instructions is within the scope of the action of the same barrier instruction, it is considered that the barrier instruction may affect its dependency relationship; otherwise, record it as an unprotected pair of cross-instructions.
[0079] Mark the identified pairs of cross-access instructions related to time and record their time dependency relationship.
[0080] Combine the pairs of cross-access instructions and the pairs of instructions with a time-dependent relationship, and generate a group of access instructions at risk of instruction rearrangement according to the core number, access target address, and time order:
[0081] For pairs of cross-access instructions for the same target address, check whether there is a time dependency relationship. If both the cross-relationship and the time dependency conditions are met, include this pair of instructions in the group of access instructions at risk of rearrangement.
[0082] Sort the instructions in the group of access instructions at risk according to the core number, target address, and time order. After sorting, group them by target address, and each group represents an independent set of access instructions at risk.
[0083] Store the generated instruction reordering risk access groups in the form of structured data, and the specific storage format includes but is not limited to JSON files or relational database tables. The stored data should include information such as instruction pair numbers, core numbers, target addresses, time differences, barrier associations, etc., to ensure the integrity of the data and facilitate subsequent analysis and use.
[0084] Conduct concurrent consistency analysis on the access instruction reordering risk access groups, identify race conditions that may occur during the use of low-level hardware primitives, and record the synchronization methods and memory barrier instruction distributions corresponding to the access instruction sequences, specifically including:
[0085] Classify the access instructions in the access instruction reordering risk access groups. According to the operation types, target addresses, and occurrence times of the access instructions, the access instructions in the access instruction reordering risk access groups are divided into read access instruction groups, write access instruction groups, and mixed operation access instruction groups:
[0086] The operation type of each access instruction can be divided into read operations, write operations, and mixed operations. For example, if an access instruction is to read the value of a shared variable, it is classified as a read operation; if it is to write a new value to a shared variable, it is classified as a write operation; if it includes both reading and writing, it is a mixed operation.
[0087] Extract the target addresses from the access instructions to identify the memory locations involved in each instruction. The target addresses are used to determine whether the access instructions operate on the same shared variable during the classification process.
[0088] After sorting the access instructions in chronological order, group them according to the operation type and target address. The grouping results form three main sets: read access instruction groups, write access instruction groups, and mixed operation access instruction groups.
[0089] For the access instructions in each group, analyze the usage of low-level hardware primitives, including marking the dependency relationships, scopes of action, and execution orders of the hardware access instructions:
[0090] Dependency relationship: Determine whether the access instructions depend on specific hardware primitives (such as load and store instructions), and record the call order and dependency chain between the hardware primitives. For example, some access instructions may depend on load access instructions to complete data acquisition, or depend on store access instructions to complete write operations.
[0091] Scope of action: Determine the scope of influence of the hardware primitives relied on by each access instruction on the shared variables. The scope of action includes the target address range, time range, and the coverage of the primitive on other instructions. For example, some hardware primitives may only be valid for specific shared variable addresses and will not affect data access at other addresses.
[0092] Execution order: Analyze the specific positions of each access instruction and the hardware primitives it depends on in the execution sequence, and record their execution order to ensure that the execution logic of the instruction stream can be accurately reproduced in subsequent consistency analysis.
[0093] On a grouped basis, according to the target address and time relationship between access instructions, analyze whether there are potential race conditions for access instructions on the same target address, including the situation where multiple cores access the same target address simultaneously and the memory barrier instructions fail to cover all access instructions:
[0094] Target address conflict: If multiple cores execute access instructions (such as reading and writing or writing and writing simultaneously) on the shared variable of the same target address, a race condition may occur. The identification of the race condition is judged by the overlap of the access times to the target address.
[0095] Memory barrier instruction coverage analysis: Judge whether the access instructions are covered by memory barrier instructions. For example, some memory barrier instructions may only act on part of the target addresses, resulting in some access instructions not being correctly synchronized, thus generating potential race conditions.
[0096] Record the synchronization methods involved in the access instruction sequence of each group, specifically including the type of synchronization method, the synchronization scope, and the specific access instructions it acts on:
[0097] Synchronization method type: Record the synchronization methods used in each access instruction group (such as lock mechanism, barrier instructions, etc.).
[0098] Synchronization scope: Describe the scope of action of the synchronization method, including the covered target addresses and time periods.
[0099] Associated access instructions: Indicate the association between the synchronization method and the specific access instructions. For example, some synchronization methods only protect part of the access instructions and do not fully cover the entire group.
[0100] Record the distribution of memory barrier instructions in the access instruction sequence, including the position of each memory barrier instruction and the range of access instructions it affects:
[0101] Position marking: Mark the specific position of each memory barrier instruction in the access instruction sequence.
[0102] Influence range: Describe the range of access instructions affected by each memory barrier instruction and the range of access instructions not covered.
[0103] For example, "Barrier 1" is located after the third access instruction. The "influence range" describes the access instruction numbers covered by the barrier instruction. For example, "Barrier 1" covers the first to the third access instructions. The "uncovered range" records the access instructions not protected by the barrier instruction. For example, "Barrier 1" does not cover the fourth to the fifth access instructions. Through these records, ensure that the distribution of memory barrier instructions is clearly traceable in subsequent analyses.
[0104] Inject a high-intensity concurrent operation scenario in the actual simulation environment. Use the concurrent monitoring module to continuously track the read and write operations of shared variables, and compare the execution order of each instruction reordering risk access group with the expected order obtained by parsing, specifically including:
[0105] Build an actual simulation environment, set up a multi-core high-concurrency operation scenario, including defining the target address range of shared variables and allocating access operation tasks for each core, ensuring that multiple cores perform read and write operations simultaneously:
[0106] According to the requirements of the actual application in the simulation environment to be simulated, clarify the target address range of shared variables in memory. For example, limit the shared variables to a continuous memory block to facilitate subsequent tracking and comparison of instruction operations. For example, assume that the shared variables are stored in the address range "0x1000–0x1FFF", then this range needs to be used as the shared memory area in the simulation environment.
[0107] Allocate specific access operation tasks for each core to ensure a certain degree of competition among cores for shared variables. For example:
[0108] Core 1 performs frequent read operations on shared variables; Core 2 performs write operations on shared variables; Core 3 performs both read and write operations simultaneously.
[0109] Call the multi-core scheduling function of the processor to start multiple cores running in parallel, enabling them to perform operations on shared variables within the same time period, simulating typical behaviors in a high-concurrency scenario.
[0110] Inject a high-intensity concurrent load in the simulation scenario, set each core to access shared variables at randomized time intervals to increase the competition among access instructions:
[0111] Control the access of each core to shared variables through randomly generated time intervals. For example, each core randomly performs read or write operations at different time points to disrupt the order of instructions, thus creating potential conflicts. For example: Core 1 reads the shared variable after a time interval of 100 nanoseconds; Core 2 writes to the shared variable after a time interval of 50 nanoseconds; Core 3 reads and writes to the shared variable after a time interval of 150 nanoseconds.
[0112] According to the concurrency requirements of the target application scenario, gradually increase the access frequency of the cores to the shared variables, simulating the change process from light load to high load. For example, the access frequency of each core can be gradually increased from 100 times per second to 1000 times per second.
[0113] During the process of injecting high-intensity load, record the load change situation of each core, including the time interval and access frequency, for subsequent analysis of the sequential relationship between instructions.
[0114] Use the concurrency monitoring module to continuously track all read and write operations of the shared variables, record the operation type, target address, and execution time of each access, and generate a complete access instruction sequence:
[0115] For each access of the core to the shared variable, the concurrency monitoring module records the operation type, including reading and writing. For example: At time T1, core 1 reads variable X; at time T2, core 2 writes variable X.
[0116] The concurrency monitoring module also records the target address of each access operation to ensure that the location of the shared variable in memory can be accurately located. For example: At time T1, core 1, address "0x1010"; at time T2, core 2, address "0x1010".
[0117] Add an accurate timestamp to each access operation to record its occurrence time. For example: Time T1 = 10 nanoseconds; Time T2 = 15 nanoseconds.
[0118] Combine the recorded operation type, target address, and execution time to form a complete access instruction sequence. For example:
[0119] Access instruction sequence of core 1: [Read 0x1010 T1], [Write 0x1020 T2]; Access instruction sequence of core 2: [Read 0x1010 T3], [Write 0x1020 T4].
[0120] Compare the tracked access instruction sequence with the expected order obtained by parsing, and analyze whether there is instruction reordering in the actual execution order:
[0121] Import the expected order of access instructions from the parsing module, and the expected order is sorted according to strict address and time dependencies. For example: Expected order: [Read 0x1010 Core 1 T1] → [Write 0x1010 Core 2 T2].
[0122] Import the actual execution order of access instructions from the concurrency monitoring module and compare it item by item with the expected order. For example: Actual order: [Write 0x1010 Core 2 T2] → [Read 0x1010 Core 1 T3].
[0123] If the actual order is inconsistent with the expected order, it is recorded as an instruction reordering phenomenon. For example, in the above example, the actual order is reordered because the write operation of Core 2 is executed earlier than the read operation of Core 1.
[0124] Store the detected instruction reordering situations in the comparison results in the form of structured data, recording the cause, scope of influence, and the access instruction numbers involved in each instruction reordering:
[0125] Record the cause of each instruction reordering phenomenon. For example, due to the incorrect effectiveness of the memory barrier instruction, the access instruction order of the two cores conflicts.
[0126] Describe the scope of influence of the instruction reordering on the shared variables, including the target addresses affected and the specific numbers of the instructions. For example: Address range: [0x1010]; Affected instructions: [Read Core 1 T1], [Write Core 2 T2].
[0127] Store the reordered data in the form of structured data, such as JSON or CSV format, to ensure convenient retrieval and use in subsequent analysis.
[0128] Statistically analyze the execution sequence matching degree of each instruction reordering risk access group under different simulation load conditions, and comprehensively generate real-time simulation performance evaluation indicators for the weak memory model conditions, specifically including:
[0129] Set multiple load conditions in the simulation environment. The load conditions include high-intensity load, medium-intensity load, and low-intensity load. Under each load condition, different concurrent scenarios are formed by adjusting the access frequency and operation interval of multiple cores to shared variables:
[0130] High-intensity load: Under high-intensity load conditions, each core frequently accesses the shared variables at the minimum time interval, simulating an extremely high-concurrency environment. For example, set the access frequency of each core to 1000 times per second, and the operation interval for each access is less than 10 microseconds.
[0131] Medium-intensity load: Under medium-intensity load conditions, each core accesses the shared variables at a relatively balanced time interval. For example, the access frequency of each core is 500 times per second, and the operation interval is between 10 microseconds and 50 microseconds.
[0132] Low-intensity load: Under low-intensity load conditions, each core accesses the shared variables at a relatively large time interval, simulating a low-concurrency environment. For example, set the access frequency of each core to 100 times per second, and the operation interval is greater than 50 microseconds.
[0133] Under each load condition, the access time interval of the core to the shared variable is adjusted by introducing a randomization mechanism. For example, the access order and time interval between cores can be generated by a random function, so as to more realistically simulate the non-deterministic operation mode of multi-cores.
[0134] For each instruction reordering risk access group under each load condition, compare each group one by one according to the actual execution order and the expected order of the access instruction sequence, and record the execution order matching degree of each group of access instructions. The execution order matching degree is determined by comparing whether the actual execution order is consistent with the expected order obtained by parsing:
[0135] From the access instruction sequence collected from the simulation environment, obtain the actual execution order of the instruction reordering risk access group. For example: The actual order of instruction group 1 is: [Read 0x1000, time 10 microseconds] → [Write 0x1000, time 15 microseconds]; The actual order of instruction group 2 is: [Write 0x2000, time 20 microseconds] → [Read 0x2000, time 25 microseconds].
[0136] Define the access order of each instruction group according to the expected order generated by the foregoing parsing steps. For example: The expected order of instruction group 1 is: [Read 0x1000, time 10 microseconds] → [Write 0x1000, time 15 microseconds]; The expected order of instruction group 2 is: [Read 0x2000, time 20 microseconds] → [Write 0x2000, time 25 microseconds].
[0137] The execution order matching degree is calculated by the following formula: ; where represents the execution order matching degree, represents the number of instruction pairs that match the expected order in the actual order, represents the total number of instruction pairs.
[0138] For each instruction reordering risk access group, calculate the matching degree according to the difference between the actual execution order and the expected order. For example: There are 2 instruction pairs in instruction group 1, and the actual order is exactly the same as the expected order, so the matching degree is 1; There are 2 instruction pairs in instruction group 2, but only 1 matches the expected order, so the matching degree is 0.5.
[0139] Calculate the average execution order matching degree of each instruction reordering risk access group under each load condition. The matching degree data includes the statistical results of perfect match, imperfect match and non-match:
[0140] The matching degree data of each instruction group refers to Table 1 Instruction Group Matching Degree Table:
[0141] Table 1 Instruction Group Matching Degree Table
[0142]
[0143] The average matching degree is the ratio of the sum of the execution order matching degrees of all instruction groups to the total number of instruction groups.
[0144] Based on the matching degree data under various load conditions, generate real-time simulation performance evaluation metrics for the weak memory model. The real-time simulation performance evaluation metrics include the execution consistency of each instruction reordering risk access group under different load conditions, the frequency of instruction reordering, and the instruction execution time deviation:
[0145] Execution consistency:
[0146] Statistically calculate the average matching degree of each instruction group under different load conditions. For example: the average matching degree under high-intensity load is 0.75; the average matching degree under medium-intensity load is 0.8; the average matching degree under low-intensity load is 1.0.
[0147] Statistically calculate the total number of times of instruction reordering and its proportion under all load conditions. For example: in high-intensity load, instruction reordering occurs 10 times, accounting for 50%; in medium-intensity load, instruction reordering occurs 5 times, accounting for 25%.
[0148] Calculate the deviation between the actual execution time and the expected execution time: ; where represents the time deviation, represents the actual execution time, represents the expected execution time.
[0149] It should be noted that the present invention is designed for the ARMv8 architecture. Its applicability stems from the default weak memory model of ARMv8, which allows out-of-order execution and memory read / write reordering to optimize high-performance and low-power applications. However, this weak sequential consistency is more likely to cause instruction reordering and consistency problems in a high-concurrency environment, so a special evaluation method is required. In contrast, other ARM architectures such as ARMv7 more adopt strong consistency strategies, and the impact of out-of-order execution and memory reordering is smaller, and the same degree of synchronization problems as ARMv8 will not occur. The solution of the present invention is deeply customized according to the characteristics of the ARMv8 architecture and can effectively cope with the simulation consistency challenges brought by its unique memory model.
[0150] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0151] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0152] Those of ordinary skill in the art can realize that the modules and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0153] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0154] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in an electrical, mechanical, or other form.
[0155] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical module, and it may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0156] In addition, in each embodiment of this application, each functional module can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0157] If the above function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0158] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0159] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for evaluating the real-time simulation performance of a multi-core processor based on the ARMv8 architecture, characterized in that, The steps are as follows: S1: Set up a concurrent monitoring module in the multi-core processor to collect the access instruction sequences of each processing core to shared variables and the corresponding memory barrier instruction information; S2: Analyze the collected access instruction sequences and memory barrier instruction information, and generate instruction reordering risk access groups based on the address cross-relationship and timing dependency between access instructions, specifically including: Analyze the collected access instruction sequences, extract the target address, operation type, and occurrence time of each access instruction, and sort the access instructions of the same core according to the occurrence time to form an in-core access sequence; Based on the in-core access sequence obtained by parsing, identify the cross-relationship of target addresses, determine whether multiple cores access the memory address of the same shared variable simultaneously, and record the cross-access instruction pairs; Analyze the time information of the cross-access instruction pairs, and judge whether there is a timing dependency according to the occurrence time of the access instructions and their correlation with the memory barrier instructions, and record the instruction pairs with a time-dependent relationship; Combine the cross-access instruction pairs and the instruction pairs with a time-dependent relationship, and generate instruction reordering risk access groups according to the core number, access target address, and time order; S3: Conduct a concurrent consistency analysis on the instruction reordering risk access groups, identify the race conditions that may occur in the use of low-level hardware primitives, and record the synchronization method and memory barrier instruction distribution corresponding to the access instruction sequences; S4: Inject a high-intensity concurrent operation scenario in the actual simulation environment, use the concurrent monitoring module to continuously track the read and write operations of shared variables, and compare the execution order of each instruction reordering risk access group with the expected order obtained by parsing; S5: Statistically analyze the execution sequence matching degree of each instruction reordering risk access group under different simulation load conditions, and comprehensively generate real-time simulation performance evaluation indicators for the weak memory model condition.
2. The real-time simulation performance evaluation method based on the ARMv8 architecture multi-core processor according to claim 1, wherein Set up a concurrent monitoring module in the multi-core processor to collect the access instruction sequences of each processing core to shared variables and the corresponding memory barrier instruction information, specifically including: Configure the concurrent monitoring module to record the access operations of each core in the multi-core processor to shared variables, collect access instruction information, and the access instruction information includes the operation type, target address, and its occurrence time; Capture memory barrier instructions in real time in the concurrent monitoring module, and extract the type, specific insertion position of the memory barrier instructions, and their correlation with access instructions; Summarize the access instruction information and memory barrier instruction information collected by the concurrent monitoring module, and perform time correction on the collected data based on a unified time reference to generate a global record including the access instruction sequences of all processing cores and memory barrier instruction information.
3. The real-time simulation performance evaluation method based on the ARMv8 architecture multi-core processor according to claim 1, wherein Conduct a concurrent consistency analysis on the instruction reordering risk access groups, identify the race conditions that may occur in the use of low-level hardware primitives, and record the synchronization method and memory barrier instruction distribution corresponding to the access instruction sequences, specifically including: Classify the access instructions in the access instruction reordering risk access group according to the operation type, target address, and occurrence time of the access instructions. The access instructions in the access instruction reordering risk access group are divided into a read access instruction group, a write access instruction group, and a mixed operation access instruction group; For the access instructions in each group, analyze the usage of low-level hardware primitives, including marking the dependency relationships, scope of action, and execution order of hardware access instructions; Based on the grouping, analyze whether there are potential race conditions for the access instructions on the same target address according to the target address and time relationships between the access instructions, including the situation where multiple cores access the same target address simultaneously and the memory barrier instructions do not cover all access instructions; Record the synchronization methods involved in the access instruction sequence of each group, specifically including the type of synchronization method, the synchronization scope, and the specific access instructions it acts on.
4. The real-time simulation performance evaluation method based on the ARMv8 architecture multi-core processor according to claim 1, wherein Inject a high-intensity concurrent operation scenario in the actual simulation environment, use the concurrent monitoring module to continuously track the read and write operations of shared variables, and compare the execution order of each instruction reordering risk access group with the expected order obtained by parsing, specifically including: Construct an actual simulation environment, set a multi-core high-concurrency operation scenario, including defining the target address range of shared variables and allocating the access operation tasks of each core to ensure that multiple cores perform read and write operations simultaneously; Inject a high-intensity concurrent load in the simulation scenario, and set each core to access the shared variable at a randomized time interval; Use the concurrent monitoring module to continuously track all read and write operations of the shared variable, record the operation type, target address, and execution time of each access, and generate a complete access instruction sequence; Compare the access instruction sequence recorded by the tracking with the expected order obtained by parsing, and analyze whether there is an instruction reordering phenomenon in the actual execution order; Store the detected instruction reordering situation in the comparison result in the form of structured data, and record the cause, scope of influence, and access instruction numbers involved in each instruction reordering.
5. The real-time simulation performance evaluation method based on the ARMv8 architecture multi-core processor according to claim 1, characterized in that Statistically analyze the execution sequence matching degree of each instruction reordering risk access group under different simulation load conditions, and comprehensively generate real-time simulation performance evaluation indicators for the weak memory model, specifically including: Set multiple load conditions in the simulation environment, and the load conditions include high-intensity load, medium-intensity load, and low-intensity load. Under each load condition, different concurrent scenarios are formed by adjusting the access frequency and operation interval of multiple cores to shared variables; For each instruction reordering risk access group under each load condition, compare them group by group according to the actual execution order and the expected order of the access instruction sequence, and record the execution order matching degree of each group of access instructions. The execution order matching degree is determined by comparing whether the actual execution order is consistent with the expected order obtained by parsing; Based on the matching degree data of multiple load conditions, generate real-time simulation performance evaluation indicators for the weak memory model. The real-time simulation performance evaluation indicators include the execution consistency of each instruction reordering risk access group under different load conditions, the frequency of instruction reordering, and the instruction execution time deviation.
6. The real-time simulation performance evaluation method based on the ARMv8 architecture multi-core processor according to claim 5, wherein Under high-intensity load conditions, each core frequently accesses the shared variable at the minimum time interval to simulate an extremely high-concurrency environment; under medium-intensity load conditions, each core accesses the shared variable at an equalized time interval; under low-intensity load conditions, each core accesses the shared variable at a larger time interval to simulate a low-concurrency environment; Under each load condition, the access time interval of the core to the shared variable is adjusted by introducing a randomization mechanism; the access order and time interval among the cores are generated by a random function.
Citation Information
Patent Citations
Chip multi-core processor clock precision parallel simulation system and simulation method thereof
CN101788919A
Heuristic method and device suitable for instruction rearrangement of multiple transmitting processors
CN116028127A