Multi-core-oriented high-performance parallel discrete event simulation propulsion method
By constructing a parallel simulation engine and a hybrid time management protocol, combined with a multi-process/multi-threaded hybrid platform architecture, the load balancing and event processing consistency issues of discrete event simulation in a multi-core environment are solved, achieving efficient multi-core parallel simulation and improving simulation speed and efficiency.
Patent Information
- Application Number
- CN202511045395.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-14
AI Technical Summary
In a multi-core processor environment, how can we efficiently perform discrete event simulation, improve simulation speed, ensure the sequentiality and consistency of event processing, and achieve load balancing to avoid overloading a single core?
A parallel simulation engine is built, employing a hybrid time management protocol and event management service. The time management service and event management service support the parallel simulation operation. An efficient event scheduling algorithm is designed to ensure load balancing. The parallel discrete event simulation logic clearly defines the logical process execution event processing flow. A multi-process/multi-threaded hybrid platform architecture is adopted for communication and collaboration to achieve multi-core parallelization.
It improved core utilization, reduced event processing latency, ensured timing consistency, met the stringent requirements of military simulation, and enhanced simulation efficiency.
Smart Images

Figure CN120950240A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of combat simulation and deduction technology, specifically to a simulation advancement method that incorporates parallel computing, discrete event simulation, multi-core optimization, and military system modeling. Background Technology
[0002] The computational and scheduling capabilities of a simulation engine not only affect the system's operating speed but also largely determine its scale and accuracy. Currently, simulation systems are becoming increasingly large-scale, with a growing variety and number of models. For example, in the field of military system-on-systems simulation, it is necessary to simulate and calculate the forces and entities across the entire battlefield, typically requiring more than 5,000 entities, including vehicles, ships, missiles, aircraft, and satellites. The models cover motion, sensing, communication, control, and decision-making, and their physical mechanisms involve disciplines such as force, sound, heat, electricity, and optics. This poses a significant challenge to the computational and scheduling capabilities of simulation engines, a challenge that cannot be simply addressed by improving hardware capabilities.
[0003] Constrained by Moore's Law, current computing power is mainly achieved by increasing the number of computing cores. Therefore, it is particularly important for simulation engine software to make full use of the potential of multi-core parallel computing and to design reasonable model scheduling and resource allocation strategies for the characteristics of simulation.
[0004] In summary, how to efficiently perform discrete event simulation in a multi-core processor environment, improve simulation speed, ensure the sequentiality and consistency of event processing, and achieve load balancing to avoid overloading a single core are currently pressing technical challenges that need to be addressed. Summary of the Invention
[0005] This invention provides a high-performance parallel discrete event simulation advancement method for multi-core processors, which can efficiently perform discrete event simulation in a multi-core processor environment, improve simulation speed, ensure the sequentiality and consistency of event processing, and achieve load balancing to avoid overloading of a single core.
[0006] This invention provides a high-performance parallel discrete event simulation advancement method for multi-core processors, comprising: Construct a parallel simulation engine, support parallel simulation operation through time management service and event management service, and based on the parallel simulation engine, determine the parallel discrete event simulation logic and clarify the processing flow of logical process execution events; Based on a parallel simulation engine, multi-core parallel simulation is performed, wherein the multi-core parallel simulation includes basic services, parallel discrete event simulation services, and optimization and extension services.
[0007] In some instances, the processing flow of the logical process execution events includes four levels of parallelism from top to bottom: the first level is parallel operation of multiple samples of the simulation system; the second level is parallel operation of multiple MPI processes; the third level is parallel operation of logical processes; and the fourth level is parallel operation of complex model solving.
[0008] In some instances, the time management service refers to the use of a hybrid time management protocol to divide the time management service into a global control mechanism and a local control mechanism. The local control mechanism is responsible for logical process scheduling, while the global control mechanism is responsible for global time synchronization. Logical process scheduling refers to the reasonable allocation of CPU among logical processes, while global time synchronization is the calculation of the smallest timestamp of all logical processes that may be processed in the future.
[0009] In some instances, the minimum timestamp is obtained by calculating the minimum transmission timestamp EETS, which includes: If thread core Calculation required Then submit ,renew ; When thread core i If the target thread core sends event e j It has been submitted. and Then update ; Last commit thread core i Responsible for calculation , ; If multiple processes are performing the simulation, the process with thread 0 as the last committer is determined. ; in, This represents the thread timestamp of thread core i, used to record the minimum event timestamp value in the current thread core. TKFEL is the set of FELs for all logical threads on the thread core, and EB is the buffer. This is a local variable representing thread core i, used to record the minimum timestamp of all events sent to the computation-state thread core. Indicates thread core i The earliest sent event timestamp, This represents the timestamp of the earliest event sent by the process core, as calculated. This indicates that the wall clock time is wt, and the earliest sent event timestamp is obtained through a snapshot.
[0010] In some instances, the event management service provides services including event creation, delivery, and submission through an event manager with a thread-local, double-column, multi-level stack structure. It also employs an event caching mechanism to support pointer-based asynchronous communication, uses a ring queue structure to cache events, separates event sending and receiving operations, and enables high-speed communication between thread cores. The event manager uses a lock-free creation and asynchronous submission mechanism to decouple the relationship between threads.
[0011] In some instances, the basic services include: building the basic architecture of a parallel system to support message interaction among a group of entities; the parallel discrete event simulation service includes building a parallel discrete event simulation framework; and the optimization and extension services include providing performance optimizations tailored to application characteristics.
[0012] In some instances, the parallel simulation kernel model adopts a hybrid multi-process / multi-thread platform architecture. Communication and collaboration between computing nodes and simulation kernels are carried out in a multi-process manner, while communication within computing nodes is optimized in a multi-threaded manner, and multi-core parallelization is transparently implemented.
[0013] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: (1) Supports single-node multi-core parallelism, significantly improving core utilization; (2) High event processing throughput, greatly reducing latency; (3) The timing consistency error rate is extremely low, meeting the stringent requirements of military simulation. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart illustrating the implementation of a high-performance parallel discrete event simulation method for multi-core processors, provided by an embodiment of the present invention. Figure 2 This is a schematic diagram of the computing protocol provided in an embodiment of the present invention; Figure 3 This is the processing flow for the execution of a logical process provided in the embodiments of the present invention; Figure 4 This is the logical process execution event processing flow provided in the embodiments of the present invention; Figure 5 This is the hierarchical parallel simulation kernel model architecture provided in the embodiments of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] In the following description, specific embodiments of the invention will be illustrated with reference to steps and symbols performed by one or more computers, unless otherwise stated. Therefore, these steps and operations will be referred to several times as being performed by a computer, and computer execution as referred to herein includes operations by a computer processing unit representing electronic signals of data in a structured format. This operation transforms the data or maintains it at a location in the computer's memory system, which can be reconfigured or otherwise alter the operation of the computer in a manner well known to those skilled in the art. The data structure maintained by the data is the physical location of the memory, which has specific characteristics defined by the data format. However, the principles of the invention described above are not intended to be limiting, and those skilled in the art will understand that many of the following steps and operations can also be implemented in hardware.
[0018] The terms "module" or "unit" as used herein can be considered as software objects executing on the computing system. Different components, modules, engines, and services described herein can be considered as implementations on the computing system. The apparatus and methods described herein are preferably implemented in software, but can also be implemented in hardware, both of which are within the scope of this invention.
[0019] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0020] This invention provides a high-performance parallel discrete event simulation advancement method for multi-core processors. By using multi-core parallel processing, the advancement mechanism of discrete event simulation is optimized, reducing simulation time; an efficient event scheduling algorithm is designed to ensure the sequentiality and consistency of event processing; load balancing is achieved in a multi-core environment to avoid overloading of a single core and improve overall simulation efficiency; and an advanced synchronization mechanism is used to ensure data consistency and accuracy of simulation results among multiple cores.
[0021] like Figure 1 As shown, this paper proposes an implementation process for a high-performance parallel discrete event simulation method for multi-core processors. The implementation process first requires building a parallel simulation engine, which mainly supports parallel simulation operation through time management and event management services. Based on the parallel simulation engine, the parallel discrete event simulation logic is studied, and the processing flow of the logic process execution is clarified. The specific implementation methods of each step are as follows: 1. Build a parallel simulation engine The high-performance parallel discrete event simulation advancement method for multi-core processors relies on the operation of a parallel simulation engine, which supports time management services and event management services.
[0022] (1) Time management services Time management is a core service in parallel discrete event simulation. In the multi-core era, processor parallelism will continue to increase, but the memory required per processing core may decrease. Adopting a hybrid optimistic and conservative time management protocol can leverage the advantages of high parallelism in optimistic protocols and low memory requirements in conservative protocols. Therefore, this system uses a hybrid time management protocol, dividing the time management service into two parts: a global control mechanism and a local control mechanism. The local control mechanism is responsible for logical process scheduling, while the global control mechanism is responsible for global time synchronization.
[0023] Logical process scheduling refers to the rational allocation of CPU resources among logical processes to ensure the rapid and accurate progress of simulation. At the beginning of each loop, the thread core needs to select a logical process, set a time frame for its execution, and then transfer CPU execution to the logical process to advance the simulation. The scheduling algorithm used can be divided into the following two phases: (1) Conservative Phase: There are conservative logical processes that can be advanced within the thread core. The conservative logical process with the smallest timestamp for the next event is selected from these processes and advanced beyond the set time limit. (2) Optimistic phase: There is no conservative logical process that can be advanced in the thread core. Select the optimistic logical process with the smallest timestamp of the next event and advance it to after the set time limit.
[0024] This scheduling method operates on a process-by-process basis, maximizing the scope of each logical process's advancement to reduce the number of logical process switches and thus decrease overhead. It's important to note that the scheduling algorithm mentions two time limits: the former is the upper limit of the timestamp that the current conservative logical process can advance, closely related to the global time synchronization discussed below; the latter can be set according to the specific application, currently defaulting to the minimum next event timestamp of other logical processes within the thread core.
[0025] The specific task of global time synchronization is to calculate the minimum timestamp of future events that all logical processes can handle, denoted here as (EIT). For conservative logical processes, this value is used as LBTS, defining the upper limit of the timestamp for events that can be safely executed (without causing a rollback); while for optimistic logical processes, this value is used as GVT, which is the lower limit of the timestamp for submitting events, completing memory release, and I / O interaction. When calculating EIT, existing LBTS barrier algorithms and GVT algorithms can be used. However, regardless of the algorithm, the processes participating in the simulation need to submit their minimum send timestamp (EETS). In a multi-process simulation kernel model, each process has only one scheduling center, making it very simple to determine EETS. However, in a hierarchical parallel simulation kernel model, there are multiple scheduling centers, and calculating the EETS value without interfering with the normal progress of thread kernels is relatively complex. Superficially, this problem can be solved by porting the GVT algorithm based on the shared memory model. The algorithm controls synchronization through a global variable GVTFlag, with each processor asynchronously detecting the GVTFlag status and participating in the calculation of the synchronization value. However, the algorithm is geared towards a purely optimistic mode and does not support mixed-time progression modes; moreover, to optimize simulation efficiency, global time synchronization calculations are typically required to be initiated at different times, but a single global variable is insufficient to meet diverse synchronization needs; thirdly, the algorithm does not consider the impact of inter-node communication on global synchronization values. To address these issues, this system proposes a computation protocol that can be flexibly configured as the EETS algorithm, supports thread cores in mixed states, and can compute EETS without blocking the normal progress of thread cores.
[0026] Variable definition : The thread timestamp of thread core i, used to record the minimum value of event timestamps in the current thread core; TKFEL is the set of FELs of all logical threads on the thread core; EB is the buffer (including events and anti-events); : A local variable of thread kernel i, used to record the minimum timestamp of all events sent to the computational thread kernel; : The earliest sent event timestamp of thread core i; The earliest sent event timestamp of the process core, calculated by the algorithm; The wall clock time is wt, and the earliest sent event timestamp (the actual EETS) is obtained through a system snapshot.
[0027] Thread core computation protocol 1) If thread core i needs to calculate EETS, submit ,renew ; 2) When thread core i sends event e, if the target thread core j has already committed... and Then update ; 3) The last commit The thread core i is responsible for computation , ; 4) If multiple processes are performing the simulation, the last one to submit is the one from thread 0. .
[0028] As long as all thread cores in the process are committed. This allows for the acquisition of an acceptable EETS value. The protocol provides a flexible framework for EETS calculation, enabling each thread core to rationally choose when to participate in EETS calculation based on real-time requirements. Assuming the thread core is at wall clock time... Submit sequentially The final process count is then determined by the last thread core to submit the data. Value. Thread core submission. The wall clock time constitutes a truncation. Before truncation is the preparatory state, and after truncation is the calculation state. Therefore, the sender and receiver states of the event can be divided into four categories: Preparation to calculation; : From preparation to preparation; : Calculation to preparation; : Calculation to calculation.
[0029] Theorem 1: If all thread cores participate in the EETS calculation, and no messages are received from other processes during the calculation, then the EETS value calculated according to basic algorithm rules 1 to 3 satisfies... .
[0030] prove: (1) First prove
[0031] Assumption ∈ It is to satisfy The event j is satisfied The thread core, then .
[0032] (1.1) If Let there exist k such that Otherwise, then it exists. ,satisfy .
[0033] We can analyze it in two scenarios: Case I: There is an ancestor event , .according to Definition, .so ,and contradiction; Situation II: There is no belonging to Ancestor events Then there must be an ancestral event. , or .so ,according to The definition has Therefore, based on the assumption in case (1.1), we have Contradictory, therefore, if There exists a k such that... ; (1.2) If Let there exist k such that .
[0034] • if The conclusion is obviously true; • like Then there exists an event. ,satisfy In thread core j submission Received it later.
[0035] if ,but ,and contradiction; if This can also be analyzed in two cases: Case I. There is an ancestor event , ,and contradiction; Case II. There is no belonging to If there is an ancestor event, then there must be an ancestor event. , or ,so .according to The definition has ,and contradiction.
[0036] Therefore, if There exists a k such that .
[0037] Situation (1.1) and Situation (1.2) combined:
[0038] (2) Then prove
[0039] By definition, it is obvious that for any i, .
[0040] Take the definition from the previous text If generated The wall clock time is greater than , there must be If generated The wall clock time is less than But because of sending The wall clock time is greater than ,so Father's incident exist It is currently being executed (i.e., not yet completed).
[0041] so, ,have .
[0042] In summary, .
[0043] If a remote message is received from another process during the calculation of EETS, let's assume... In this type of event, there is a minimum timestamp if Then Theorem 1 still holds. And if Then we have: Theorem 2. If all thread cores participate in the EETS computation, and Then, the EETS values calculated according to basic algorithm rules 1 to 4 satisfy... .
[0044] Proof: According to The definition, and because Obviously there is and According to computation protocol 4, thread core 0 submits tEETS last, at which point the wall clock is... It may be forwarded by the corresponding C logic process, or cached in the TKFEL of the C logic process. Let's assume... ,So: • If it is forwarded, then : Case I: According to rule 2, there is ,so ,So ; Situation II: : like Executed. ,so ; like Not executed. Similarly, g; therefore .
[0045] • If cached in TKFEL, press Definition, have .so .In summary, .
[0046] According to Theorem 1 and Theorem 2, It is confined to a reasonable range. Considering... It was used to calculate EIT, and Rule 4 is necessary and sufficient. The computation protocol defines a flexible framework that allows users to specify which thread cores participate in the EETS computation based on the characteristics of the simulation application, in order to configure the required global synchronization algorithm, such as... Figure 2 As shown. If the risk of optimistic execution in the simulation application is high, the thread kernel can commit tEETS immediately after entering optimistic mode; if the risk of optimistic execution is low, the thread kernel's optimistic execution is allowed to advance to a certain timestamp limit before committing tEETS. The function NeedUpdateEETS() provides a configurable interface for users to determine whether the thread kernel needs to commit tEETS. However, since the hierarchical parallel simulation kernel model uses a scheduling strategy based on logical processes, the interval between two calls to NeedUpdateEETS() may be long, thus delaying EETS calculation and EIT calculation. After the thread kernel sends an event, it can immediately participate in EETS calculation when it detects that the target thread kernel has committed. However, if the thread kernel does not send messages for a long time, it cannot detect whether it can commit tEETS, which is where NeedUpdateEETS() comes in. The two detection mechanisms work together to control the thread kernel's participation in EETS calculation at different granularities, thereby achieving efficient global synchronization.
[0047] (2) Event Management Service Event management service is a fundamental service, primarily providing simulation systems with services such as event creation, transmission, and submission. In parallel discrete event simulations, there are typically numerous event interactions between logical processes, and each event must undergo a creation, transmission, and submission process. The efficiency of the event management service has a significant impact on the overall performance of the simulation platform.
[0048] Event creation and commit are a pair of dual operations, corresponding to the `new` and `delete` operations on memory, respectively. While general-purpose memory allocators exist to support multithreaded applications, due to the unique system architecture and application characteristics of hierarchical parallel simulation kernel models, targeted design of the event allocation and deallocation mechanism helps achieve optimal performance. (1) A general-purpose memory allocator requires a complex mechanism to support applications with dynamically changing thread counts. In contrast, in a hierarchical parallel simulation kernel model, the number of threads remains fixed after the simulation runs. (1) The general memory allocator aims to minimize the latency of both new and delete. However, in the hierarchical parallel simulation kernel model, the address pointer needs to be returned as soon as possible after the event is created, which has high real-time requirements; while the real-time requirements of event submission are relatively low. The number of event types is much smaller than the number of events; Events within and between thread kernels are executed by different threads. Physically separating memory can reduce cache false sharing conflicts.
[0049] To address these characteristics, this system designs an event manager with a thread-local, dual-column, multi-level stack structure. It employs a lock-free creation and asynchronous commit mechanism to decouple the relationship between threads, thereby achieving efficient event services.
[0050] In hierarchical parallel simulation kernel models, there are three types of communication: intra-thread core, inter-thread core, and inter-process. Intra-thread core communication is straightforward; messages are simply inserted into the future event queue (FEL) of the target logical process. Inter-process communication equals two inter-thread core communications plus one network communication, but multi-core processors cannot improve network communication. Therefore, optimization primarily focuses on inter-thread core communication. Since all threads within a process can share the address space, pointers can be used for event passing. This significantly improves communication efficiency and reduces memory consumption. However, during simulation execution, threads proceed in parallel, and synchronous communication can cause mutual interference, greatly impacting simulation efficiency. This paper proposes an event caching mechanism to support pointer-based asynchronous communication and employs a ring queue structure to cache events, separating event sending and receiving operations to achieve high-speed communication between thread cores.
[0051] 2. Parallel Discrete Event Simulation A core problem to be solved by a high-performance parallel discrete event simulation advancement method for multi-core processors is how to ensure that all logical processes process events in timestamp order. Currently, there are two main types of time management protocols to address this problem: conservative protocols and optimistic protocols. Under conservative protocols, events can only be executed if the timestamp order is guaranteed not to be violated. Conservative protocols work well for simulation applications with good parallelism. However, conservative protocols are too strict, restricting the execution of certain events that do not violate timestamp order, causing processor idle waiting and thus affecting efficiency. Optimistic protocols, on the other hand, relax execution restrictions, allowing logical processes to execute events in the currently visible timestamp order. When timestamp out-of-order does occur, errors are corrected through rollback or other methods. Compared to conservative protocols, optimistic protocols are better at uncovering the potential parallelism of simulation applications, but if the rollback overhead is too high, it can also affect performance or even be counterproductive. For ease of description, the formal definitions of simulation time, simulation events (messages), and logical process paradigms are given below.
[0052] (1) Simulation time Definition 1 (Simulation Time). Simulation time T is defined by two relationships. and They are respectively: ; .
[0053] Simulation time is a fundamental concept in simulation, used to distinguish the chronological order of events. Intuitively, the first digit corresponds to the physical time of the physical system, and the second digit distinguishes events with a causal relationship within the same physical time. Unless otherwise specified, the timestamp in this system refers to the simulation time.
[0054] (2) Simulation events Definition 2 (Simulation Event). The basic unit for a logical process to execute a task. The passing of events between logical processes is called sending a message; the terms "message" and "event" will be used interchangeably below.
[0055] 1) The timestamp of the event, i.e., the simulation time at which the event was executed; 2) : Indicates the positive or negative of an event. Negative events are generated during logical process rollback and are used to cancel events caused by erroneous execution.
[0056] (3) Logical process Definition 3 (Logical Process). A logical process (Logica logical process, or simply logical process) is defined as a seven-tuple structure. : FEL: A linked list of events sorted in ascending order by timestamp, used to record unprocessed events; PEL: A linked list of events sorted in ascending order by timestamp, used to record processed events; Used to cache negative messages that arrive before the positive message; : The current state of the logical process, where ts(s) represents the simulation time of this state; A checkpoint sequence arranged in ascending order of timestamps, used to record the state of the logical process and the set of events sent at each historical point in time. Optimistic logical processes can use the checkpoint sequence to restore to the most recent correct state; for each checkpoint; : Event handling function, which modifies the state of the logical process and generates a new set of events (or can generate 0 events); : The rollback function restores the state of the logical process to before the timestamp of the input event and sends out the counter-events corresponding to all out-of-order execution events.
[0057] Figure 3 This describes the processing flow of the logical process. Each time, the next event is retrieved and processed according to the event type and timestamp. Since the conservative logical process does not encounter outdated or reversed messages, it does not need to save historical states; only lines 4 and 5 are executed. It should be noted that if the input queue is not empty, each message sent to the logical process must be matched with a reversed event in the input queue: if a match is successful, they cancel each other out; if not, the message is added to the logical process's input queue. As can be seen from the algorithm, to ensure the smooth execution of the logical process, some basic services need to be provided at the underlying level, such as `send()` for sending messages, global synchronization, etc.
[0058] Figure 4This paper proposes a logical process execution event processing flow. From the perspective of the entire system operation, parallel discrete event simulation applications exhibit four levels of parallelism from top to bottom. The first level is multi-sample level parallel operation of the simulation system: since analysis and simulation require multiple runs of multiple samples, parallel operation of multiple samples can significantly improve simulation efficiency. The second level is multi-MPI process level parallelism: logical processes can be divided into different MPI process groups according to the closeness of their interaction, which can ensure the parallel execution of different MPI processes and improve the communication efficiency between closely interacting logical processes. The third level is logical process level parallelism: each MPI process contains several logical processes, and the concurrency between these logical processes can be further developed to achieve parallel processing of multiple entity events. The fourth level is parallelism in solving complex models: in parallel simulation systems, when the model becomes complex enough or multiple targets need to be solved, using a single processing core for solving may become a performance bottleneck affecting the entire system. Parallel solving of the model is required to improve the system's operating efficiency.
[0059] 3. Parallel simulation operation for multi-core processors The multi-core parallel simulation kernel provides a series of services to support communication and collaboration between logical processes, thereby ensuring the correctness and efficiency of the entire parallel simulation. These services are divided into three categories from bottom to top: basic services, PDES services, and optimized extension services. (1) Basic services: Build the basic architecture of the parallel system to support message interaction among a group of entities (logical processes), including naming services and event management services; (2) Parallel Discrete Event Simulation Service: Build a parallel discrete event simulation framework, including time management service, rollback service, etc.; (3) Optimize and extend services: Provide performance optimizations based on application characteristics, including dynamic migration services, commonly used scientific computing algorithm services, etc.
[0060] like Figure 5As shown, to fully utilize multi-core CPU resources, multi-threading is typically used to distribute the computational load and optimize software performance. This paper proposes a hierarchical parallel simulation kernel model. This model addresses the differences in interaction capabilities between internal and external computing nodes in a multi-core cluster by employing a hybrid multi-process / multi-threaded platform architecture. Communication and collaboration between computing nodes and simulation kernels are achieved through multi-process methods; within computing nodes, multi-threading optimizes communication and transparently implements multi-core parallelism. Thus, from a system perspective, the hierarchical parallel simulation kernel model is identical to existing simulation platforms; however, from the perspective of a single node, it integrates multiple simulation scheduling and processing cores into a single process, divided into two layers: the first layer, called the Process Kernel, controls the execution of all thread kernels in the second layer, including creation, initialization, startup, and shutdown. Once started, the Process Kernel no longer occupies any computational resources, only providing global variables for synchronization. After all thread cores have finished executing, CPU execution is returned to the process core, which then handles the cleanup. The second layer consists of a group of thread cores, each of which can be considered a simplified simulation kernel responsible for scheduling logical processes, executing and sending events, etc. Each thread core corresponds one-to-one with an operating system thread, and the maximum value can be set to the number of CPU processing cores.
[0061] Figure 5This paper proposes a hierarchical parallel simulation kernel model. This model addresses the differences in interaction capabilities between internal and external computing nodes in a multi-core cluster by employing a hybrid multi-process / multi-threaded platform architecture. Communication and collaboration between computing nodes and simulation kernels are achieved through multi-process methods; within computing nodes, multi-threading optimizes communication and transparently implements multi-core parallelism. From a system perspective, the hierarchical parallel simulation kernel model is identical to existing simulation platforms. However, from the perspective of a single node, it integrates multiple simulation scheduling and processing cores into a single process, divided into two layers: the first layer, called the Process Kernel, controls the execution of all thread kernels in the second layer, including creation, initialization, startup, and shutdown. Once started, the Process Kernel no longer occupies any computing resources, only providing global variables for synchronization. After all thread kernels have finished executing, CPU execution is returned to the Process Kernel for cleanup. The second layer consists of a group of thread kernels, each of which can be considered a simplified simulation kernel responsible for scheduling logical processes, executing and sending events, etc. Each thread core corresponds one-to-one with an operating system thread, with the maximum value set to the number of CPU processing cores. To support inter-node communication, a group of Communication Logical processes (C logical processes) act as proxies to handle communication with their respective nodes. The C logical processes are called logical processes because their sent and pending event lists can be mapped to the FEL and PEL of the logical process. Event execution is defined as an event sent to the target logical process. In this way, the C logical processes are incorporated into the framework of the logical process paradigm, and inter-node message sending can be controlled using optimistic or conservative methods. This group of C logical processes is uniformly placed on thread core 0, and all inter-node messages are sent to thread core 0 for forwarding. Logically, inter-node communication is transformed into inter-thread core or intra-core communication.
[0062] The above provides a detailed description of a high-performance parallel discrete event simulation advancement method for multi-core processors, as provided by the embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A high-performance parallel discrete event simulation advancement method for multi-core processors, characterized in that, include: Construct a parallel simulation engine, support parallel simulation operation through time management service and event management service, and based on the parallel simulation engine, determine the parallel discrete event simulation logic and clarify the processing flow of logical process execution events; Based on a parallel simulation engine, multi-core parallel simulation is performed, wherein the multi-core parallel simulation includes basic services, parallel discrete event simulation services, and optimization and extension services.
2. The method according to claim 1, characterized in that, The processing flow of the logical process execution event includes four levels of parallelism from top to bottom. The first level is parallel operation of multiple samples in the simulation system; the second level is parallel operation of multiple MPI processes; the third level is parallel operation of logical processes; and the fourth level is parallel operation of complex model solving.
3. The method according to claim 2, characterized in that, The time management service refers to the use of a hybrid time management protocol to divide the time management service into a global control mechanism and a local control mechanism. The local control mechanism is responsible for logical process scheduling, while the global control mechanism is responsible for global time synchronization. Logical process scheduling refers to the reasonable allocation of CPU among logical processes, while global time synchronization is the calculation of the smallest timestamp of all logical processes that may be processed in the future.
4. The method according to claim 3, characterized in that, The minimum timestamp is obtained by calculating the minimum transmission timestamp EETS. The calculation of the minimum transmission timestamp EETS includes: If thread core Calculation required Then submit ,renew ; When thread core i If the target thread core sends event e j It has been submitted. and Then update ; Last commit thread core i Responsible for calculation , ; If multiple processes are performing the simulation, the process with thread 0 as the last committer is determined. ; in, This represents the thread timestamp of thread core i, used to record the minimum event timestamp value in the current thread core. TKFEL is the set of FELs for all logical threads on the thread core, and EB is the buffer. This is a local variable representing thread core i, used to record the minimum timestamp of all events sent to the computation-state thread core. Indicates thread core i The earliest sent event timestamp, This represents the timestamp of the earliest event sent by the process core, as calculated. This indicates that the wall clock time is wt, and the earliest sent event timestamp is obtained through a snapshot.
5. The method according to claim 4, characterized in that, The event management service provides services including event creation, delivery, and submission through an event manager with a thread-local, dual-column, multi-level stack structure. It also employs an event caching mechanism to support pointer-based asynchronous communication, uses a ring queue structure to cache events, separates event sending and receiving operations, and enables high-speed communication between thread cores. The event manager uses a lock-free creation and asynchronous submission mechanism to decouple the relationship between threads.
6. The method according to claim 5, characterized in that, The basic services include: building the basic architecture of a parallel system to support message interaction among a group of entities; the parallel discrete event simulation service includes building a parallel discrete event simulation framework; and the optimization and extension services include providing performance optimizations tailored to application characteristics.
7. The method according to claim 6, characterized in that, The parallel simulation kernel model adopts a hybrid platform architecture of multi-process / multi-thread. Communication and collaboration between computing nodes and between simulation kernels are carried out in a multi-process manner. Within the computing nodes, communication is optimized in a multi-thread manner, and multi-core parallelization is transparently implemented.