An Accelerating Method for In-Memory Computing Architecture Simulator Based on Dynamic Domain Decomposition
The mapping files generated by the compiler guide dynamic domain decomposition, which solves the problem of inefficiency of the simulator in the integrated storage and computing architecture, realizes efficient parallel processing and flexible adaptability, and improves simulation speed and performance.
Patent Information
- Application Number
- CN202510280385.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-11
AI Technical Summary
Traditional domain decomposition strategies are difficult to effectively apply to the integrated storage and computing architecture, resulting in inefficiency of the simulator. Especially in architectures containing a large number of closely connected computing cores, existing methods rely on polling or serial execution, resulting in efficiency bottlenecks.
The map file generated by the compiler guides dynamic domain decomposition, adjusts the domain division of the simulation model in real time according to task requirements, and allocates calculation threads for each part. A parallel discrete event simulation mechanism is adopted to ensure load balancing and event sequence consistency.
It significantly improves the efficiency and flexibility of the integrated storage and computing architecture simulator, can adapt to different computing tasks, maximize the utilization of hardware resources, and achieve efficient parallel processing.
Smart Images

Figure CN119781993B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent computing, and particularly relates to an acceleration method for an in-memory computing architecture simulator based on dynamic domain decomposition. Background Art
[0002] In recent years, with the expansion of the parameter scale of DNN models, devices based on CMOS and following the von Neumann architecture are now facing the challenge of the memory wall and encountering bottlenecks in terms of memory, bandwidth, and energy consumption improvement. In-memory computing is considered a promising technology to avoid the memory wall and has become a hot research direction in the field of deep neural network accelerators.
[0003] At the same time, in-memory computing simulators are playing an increasingly important role in the co-design of in-memory computing software and hardware and architecture search. To improve simulation efficiency, the decomposition method is widely used in the traditional simulation field to accelerate the simulation process, that is, the entire simulation task is divided into multiple relatively independent subtasks and run in parallel on different processor cores. This method is effective for conventional computer architectures. However, for in-memory computing architectures, their complexity lies in the fact that they contain a large number of closely connected computing cores and form a specific topology, which makes it difficult to directly apply traditional domain decomposition strategies. Therefore, in actual operation, many in-memory computing architecture simulators still rely on polling each computing core or using less efficient methods such as serial discrete event simulation, which leads to an efficiency bottleneck in the simulation process.
[0004] To address the above problems, the present invention proposes a new method based on dynamic domain decomposition, aiming to solve the efficiency problem of in-memory computing architecture simulators. This method makes full use of the mapping file information provided by the compiler corresponding to the simulation architecture, which details the task allocation of each computing core during the program execution process. Through the designed algorithm, the simulation model can be dynamically decomposed according to specific task requirements. This means that in the face of different types of computing tasks, the simulator can automatically adjust its internal structure, optimize the parallel processing ability, and thus significantly improve the simulation speed.
[0005] In addition, this method of dynamic domain decomposition not only improves the simulation efficiency but also enhances the flexibility and adaptability of the simulator. It allows the simulator to adjust its own configuration in real time according to the changes in the specific application scenario to ensure the best performance. At the same time, since this method can better explore the inherent parallelism of in-memory computing architectures, it also provides new ideas and technical support for future high-performance computing research. This invention further optimizes the in-memory computing architecture simulation technology and is expected to promote the rapid development of related fields. Summary of the Invention
[0006] The object of the present invention is to provide an acceleration method for a memory - in - computing architecture simulator based on dynamic domain decomposition aiming at the deficiencies of the prior art. The object of the present invention is achieved through the following technical solutions: An acceleration method for a memory - in - computing architecture simulator based on dynamic domain decomposition, including:
[0007] Based on a compiler, convert the target network into a form suitable for execution on a memory - in - computing architecture, that is, a binary instruction program for simulation, and additionally create a mapping file; the mapping file includes weight replication information, weight - to - memory - in - computing device array mapping information, and the number of input times corresponding to each weight block.
[0008] Obtain a configuration file, the simulator initializes and instantiates a simulation model based on the configuration file, and then performs domain decomposition on the simulation model based on the mapping file; the simulator assigns computing threads to each decomposed part, loads the binary instruction program, and the simulator starts the simulation task. The memory - in - computing architecture includes a memory - in - computing device array; the memory - in - computing device array includes several memory - in - computing cores.
[0009] Further, it also includes: measuring the performance of the simulation process based on evaluation metrics, and the evaluation metrics include latency, energy consumption, power consumption, and accuracy.
[0010] Further, the target network includes a deep neural network model.
[0011] Further, based on a compiler, convert the target network into a form suitable for execution on a memory - in - computing architecture, that is, a binary instruction program for simulation, and additionally create a mapping file, including:
[0012] First, the compiler performs operator fusion; then the weight mapping process distributes the weights of the target network into the memory - in - computing device array; the subsequent data flow scheduling focuses on optimizing the flow path of data between different computing units; finally, the compiler generates a binary instruction program for simulation and additionally creates a mapping file.
[0013] Further, the simulator assigns computing threads to each decomposed part, including:
[0014] Weight replication is used to balance the execution time between different network layers; when assigning computing threads, combined with the number of input times corresponding to each weight block, to keep the workload of each thread balanced; or assign the computing cores mapping the same network layer to the same thread; where each input of each weight block is equivalent to a unit workload of the simulator.
[0015] Further, the computing cores assigned to the same thread share a time queue for simulation.
[0016] Further, the computing cores mapped to the same network layer are assigned to the same thread. Specifically: the n network layers are respectively assigned to n threads; to ensure the execution order of instructions, the simulation process is as follows:
[0017] Thread 1 containing the computing core mapped to the first-layer weights starts simulation from time 0; when reaching time a, the number of generated results reaches the threshold, and the event queue of Thread 1 will send these data to Thread 2 containing the computing core mapped to the second-layer weights through inter-thread events; Thread 2 will start simulation from time 0 only after receiving these data, and will synchronize with the data sent by Thread 1 after its own simulation time reaches time a, and load these data into its own computing core. This can ensure that Thread 2 receives the data at the correct time; after that, Thread 2 will pause at time a until Thread 1 sends data again for simulation, and synchronize at the time of receiving new data; the interaction process between Thread n-1 and Thread n is the same as the interaction process between Thread 1 and Thread 2.
[0018] The beneficial effects of the present invention are as follows: 1. The method of the present invention is specifically designed for the in-memory computing architecture simulator, and is especially optimized for discrete event-driven simulators. This method improves the simulation efficiency and performance by modifying the existing interaction process between the compiler and the simulator, especially in the performance when dealing with complex computing tasks.
[0019] 2. The mapping file is generated by the compiler corresponding to the simulation architecture. Different from the compiled in-memory computing program, it is mainly used to describe the mapping situation of the corresponding network weights on the in-memory computing device array. It needs to include but is not limited to the following key data: weight replication information, weight and in-memory computing device array mapping information, the number of input times corresponding to each weight block, etc. These information are crucial for ensuring the accuracy and consistency of the simulation process, enabling the simulator to accurately reproduce the behavior of the original DNN model during operation. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0021] Figure 1 It is the simulation flow chart provided by the embodiment of the present invention.
[0022] Figure 2 It is the domain decomposition flow chart based on high throughput.
[0023] Figure 3 It is the domain decomposition flow chart based on low latency. Detailed implementation mode
[0024] The present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.
[0025] A method for accelerating a memory-computation integrated architecture simulator based on dynamic domain decomposition according to the present invention includes:
[0026] Based on a compiler, convert the target network into a form suitable for execution on a memory-computation integrated architecture, that is, a binary instruction program for simulation, and additionally create a mapping file; the mapping file includes weight replication information, weight and memory-computation device array mapping information, and the number of input times corresponding to each weight block;
[0027] Obtain a configuration file, the simulator initializes and instantiates a simulation model based on the configuration file, and then performs domain decomposition on the simulation model based on the mapping file. The minimum unit of domain decomposition is a memory-computation integrated computing core, that is, a functional unit that can execute instructions independently or has a complete control flow. This means that each memory-computation integrated computing core can participate in the simulation as an independent working unit. Each decomposed part can be composed of one or more memory-computation integrated computing cores, which depends on the specific computing task requirements and hardware configuration. This method ensures flexibility and scalability, and the allocation of computing resources can be adjusted according to actual needs; then, the simulator assigns computing threads to each decomposed part, loads the binary instruction program, and the simulator starts the simulation task.
[0028] Specifically, for a specific computing task, before the simulation starts, the simulator receives a network mapping file compiled by the corresponding compiler. This mapping file not only contains information about the network structure, but also details the specific location of the weight data in the memory-computation device array. Through a specific algorithm, dynamic domain decomposition is performed on the simulation model according to the mapping file to ensure that the model can be reasonably divided according to the task characteristics to adapt to different computing needs. The domain decomposition algorithm follows the principle of load balancing, striving to make the workload of each divided part in the simulator similar, so as to more effectively use multi-threading to improve the simulation speed. Through the load distribution strategy, ensure that the task distribution among various computing units is as uniform as possible, avoid the situation that some units are overloaded while other units are idle, so as to maximize the overall performance. This dynamic domain decomposition technology can effectively reduce the communication overhead between different decomposed parts and maximize the use of hardware resources.
[0029] During the simulation process, a parallel discrete event simulation mechanism is adopted. Each part generated by domain decomposition shares a global event queue for that part, and each event queue is assigned a computing thread to support efficient parallel processing while ensuring the consistency and correctness of the event order. This mechanism enables multiple computing threads to execute in parallel without conflicts, thus improving the speed and efficiency of the simulation.
[0030] The applicable architecture of this application should be a complex architecture with multiple memory - in - computing cores and a complete interconnection network, especially suitable for high - performance computing tasks in multi - core or multi - node environments, excluding simple single - core architectures. In a single - core architecture, due to the lack of parallel processing capabilities, the simulator degrades to a serial execution mode and loses the acceleration effect. Therefore, this application emphasizes the parallel acceleration advantage achieved in a multi - core environment.
[0031] This application is applicable to scenarios where the performance of the simulation host is limited. When the performance of the simulation host has no upper limit, the simulation architecture will perform domain decomposition according to the number of available memory - in - computing cores, that is, the simulation model domain will be decomposed into as many parts as there are memory - in - computing cores, and a time queue and an execution thread will be assigned to each part. In this case, under the premise of ignoring the synchronization overhead, the theoretical maximum speedup ratio can be achieved.
[0032] Figure 1 It details the workflow of the acceleration method for the memory - in - computing architecture simulator based on dynamic domain decomposition. The whole process is divided into two main stages: compilation and simulation. The compilation stage is the starting point of the whole process. On the left, a deep neural network (DNN) model is shown as the input data. The core task of this stage is to convert the original DNN model into a form suitable for efficient execution on the memory - in - computing architecture through a series of optimization steps. First, the compiler performs operator fusion, reducing the computational complexity by merging multiple operators. This step helps simplify the subsequent processing logic and improve the running efficiency. Then, the weight mapping process efficiently distributes the weights in the DNN model to the memory - in - computing device array to ensure data access locality and efficiency. Subsequently, the data - flow scheduling focuses on optimizing the data flow path between different computing units to minimize data transfer latency and power consumption, thus providing an optimal data - flow plan for the subsequent simulation. Finally, the compiler generates a binary instruction program for simulation and additionally creates a mapping file, which will guide the simulator to perform intelligent domain decomposition during the subsequent simulation process to ensure that the simulation task can effectively utilize the multi - core CPU resources.
[0033] After entering the simulation phase, the focus of work shifted to the preparation of configuration files and the execution of simulation tasks. The left side of the lower part shows the configuration files required by the simulator, which cover four key aspects: architectural resources, array parameters, device characteristics, and simulator configuration. See Table 1 for details. They provide the necessary settings for the initialization of the simulation environment. The simulator initializes and instantiates the simulation model according to these configuration files, and then intelligently partitions the simulation model based on the mapping file, that is, domain decomposition. This process ensures minimal dependencies between parts, facilitating parallel processing and improving simulation efficiency. Next, the simulator assigns computing threads to each decomposed part to ensure that each computing unit can make full use of the available resources. After loading the binary instruction program generated by the compiler, the simulator officially starts the simulation task. During this process, a series of evaluation metrics are listed on the right, including latency, energy consumption, power consumption, and accuracy. These metrics are used to comprehensively measure the performance of the simulation process and help researchers analyze and optimize the model.
[0034] Table 1. Contents of Configuration Files
[0035] This method introduces a new interaction mechanism between the compiler and the simulator - generating a mapping file through weight mapping to guide dynamic domain decomposition, and also significantly improves simulation efficiency and flexibility. Traditionally, the compiler was only responsible for generating the binary instruction program for the simulator to load. The newly added mapping file enables the simulator to perform more precise domain decomposition before simulation, further enhancing the simulator's ability to adapt to complex computing tasks. This enhanced interaction mechanism allows the simulator to more effectively utilize the multi-core CPU resources, thus improving the overall simulation speed and efficiency. Through this method, not only is the simulation speed increased, but also a new idea for in-memory computing parallel simulation is provided, which is particularly important when exploring the application scenarios of new in-memory computing architectures.
[0036] Before starting a detailed introduction to the dynamic domain decomposition algorithm and process of this application, it is necessary to explain some basic concepts in advance for better understanding. The simulation principle based on a discrete event simulator is to model and analyze the dynamic behavior of a system by processing a series of discrete events at specific time points. This method is applicable to systems where state changes are mainly triggered by discrete events, such as in the fields of computer architecture, computer networks, etc. First, it is necessary to abstract and model the target system, defining the components of the system, state variables, and the interaction rules between these components. Each component can be an entity or a resource, such as a CPU, memory, on-chip network, etc.; the state variables describe the states of these components, such as busy / idle, empty / full, etc. Next, define various events that may occur in the system, such as arrival events, service completion events, transfer events, and failure events, and set an associated timestamp for each event, indicating the specific moment when the event occurs. Create an event queue to store all future events to be processed and arrange them in chronological order. The time during the simulation does not advance continuously but jumps to the time point of the next upcoming event. After processing all events at the current time point each time, the simulator will update the system state, generate subsequent possible events based on the new state, and then add these new events to the event queue. For each event taken out from the event queue, execute the corresponding processing logic, including updating the system state, triggering new events, and recording data on relevant performance metrics. Set the conditions for ending the simulation, such as reaching a preset simulation duration, processing a certain number of tasks, or meeting a certain specific state condition. When the termination condition is met, the simulation ends and the results are output for analysis. In this way, the simulation based on a discrete event simulator can efficiently simulate the dynamic behavior of complex systems and is widely used in multiple fields such as computer architecture and computer networks, helping decision-makers understand the behavior patterns of the system, predict future trends, and formulate effective strategies.
[0037] The event queue serves as the carrier for the simulation process in a discrete event simulator. It can be considered that events in parts sharing the same event queue during the simulation process are executed strictly in chronological order. However, a single event queue can only perform serial simulation and cannot utilize the multi-core parallel characteristics of modern computers, resulting in low simulation efficiency. The basic idea of domain decomposition is to divide the simulation object into multiple parts, configure an event queue for each part, and arrange these event queues on different threads for parallel simulation. However, this brings new problems: due to different loads, the simulation speeds of different event queues may vary greatly, which may lead to inconsistent simulation times for each queue at the same moment. Event interactions will occur between different parts. If a faster queue sends an event to a slower queue, the event needs to be cached first and wait for the slower queue to execute until the interaction time to synchronize the sent event; if a slower queue sends an event to a faster queue, the situation where future events affect past states will occur on the faster queue, which may lead to serious simulation errors, generally referred to as causality errors. Therefore, the purpose of designing the dynamic domain decomposition algorithm in this application is to improve the efficiency of parallel simulation on the premise of avoiding causality errors.
[0038] Figure 2 It details the process of domain decomposition for the target memory-computation integrated architecture in the high-throughput compilation mode. The characteristics of high-throughput compilation include layer-by-layer pipelining for the inference granularity, and different batches of data are executed between layers. Each region enclosed by a dashed line in the figure represents a tile, and each geometric figure in each tile represents a memory-computation integrated computing core. In this example, the architecture has 4 tiles, and each tile has 8 memory-computation integrated computing cores. The target network is mapped to this memory-computation integrated architecture through the compiler. For ease of understanding, a diagram is used here to represent the mapping results included in the mapping file. Different geometric figures represent computing cores mapped with different network layers: circles represent the first layer (convolutional layer), squares represent the second layer (convolutional layer), diamonds represent the third layer (fully connected layer), and hexagons represent the fourth layer (fully connected layer). Different letters in the same geometric figure represent different copies of the same layer's network weights. For example, the first layer is replicated 2 times, so there are two parts, A and B; the second layer is replicated 3 times, so there are three parts, A, B, and C. The role of weight replication is to balance the execution times between different layers. By replicating multiple copies of the weights and executing them in parallel, the execution time of the layer with longer time consumption can be reduced. The lower half of the figure corresponds to the upper half in structure, except that the letters representing the replication times are replaced with numbers representing the input times. During the simulation process, each input of each piece of weight can be approximately equivalent to the unit workload of the simulator. Assuming the limit of the simulation host is 10 threads, these computing cores need to be allocated to 10 threads with the minimum sum of workloads for each thread. The bottom of the figure shows an optimal allocation result. After that, the computing cores allocated to the same thread will share the same time queue for simulation.
[0039] Figure 3 Details the process of domain decomposition for the target memory - computing integrated architecture in the low - latency compilation mode. The characteristic of the low - latency compilation mode is that for the same network inference process, when the data of the previous layer accumulates to an appropriate quantity, it will be passed to the next layer, without waiting for all the data of the previous layer to be processed. This mode is particularly suitable for application scenarios that require quick response. Although the target network and architecture are exactly the same as the examples in the high - throughput compilation, and thus the weight mapping results are the same, the data scheduling method is different. In the low - latency mode, the result of the previous layer is the input of the next layer, and there is a data dependency. If the computing cores are freely allocated as Figure 2 in [reference], it will lead to disordered instruction execution order. Therefore, the domain decomposition method adopted here is to allocate the computing cores mapping the same layer to the same thread.
[0040] As can be seen from the lower half of the figure, the four - layer network is respectively allocated to four threads. To ensure the instruction execution order, the simulation process is as follows: Thread 1 containing the computing cores mapping the first - layer weights starts the simulation from time 0. When reaching time a, the number of generated results reaches the threshold, and the event queue of Thread 1 will send this data (the output of the first - layer network) to Thread 2 containing the computing cores mapping the second - layer weights through inter - thread events. Thread 2 will start the simulation from time 0 only after receiving this data. Until its own simulation time reaches time a, it will synchronize with the data sent by Thread 1 and load this data into its own computing cores. This can ensure that Thread 2 receives the data at the correct time. After that, Thread 2 will pause at time a until Thread 1 sends data again for simulation and synchronize at the moment of receiving new data.
[0041] The interaction process between Thread 2 and Thread 3 is similar to the interaction process between Thread 1 and Thread 2. Thread 2 containing the computing cores mapping the second - layer weights starts the simulation from time 0. When reaching time b, the number of generated results reaches the threshold, and the event queue of Thread 2 will send this data (partial output of the second - layer network) to Thread 3 containing the computing cores mapping the third - layer weights through inter - thread events. Thread 3 will start the simulation from time 0 only after receiving this data. Until its own simulation time reaches time b, it will synchronize with the data sent by Thread 2 and load this data into its own computing cores. The subsequent interactions between threads also follow the same principle. Although the synchronization operation will impose additional overhead on the simulator and consume part of the advantage brought by parallel simulation, in the low - latency mode, this method can strictly ensure the instruction execution order in a parallel environment, thus ensuring the stability and accuracy of the entire system.
[0042] The method proposed in the present invention solves the above-mentioned problem by utilizing the mapping file provided by the compiler supporting the simulation architecture and combining it with a specially designed algorithm to implement dynamic domain decomposition on the simulation model. This method can adjust and optimize the division of domains in real time according to the specific needs of the simulation architecture when performing different tasks, thereby maximizing the parallel potential of the system and significantly improving the work efficiency and simulation speed of the storage-computing integrated architecture simulator. By introducing dynamic domain decomposition technology, the present invention overcomes the bottleneck in the storage-computing integrated architecture simulation, realizes a more efficient parallel simulation process, and thus improves the overall performance of the simulator.
[0043] An embodiment of the present invention also provides a storage-computing integrated architecture simulator acceleration device based on dynamic domain decomposition, including one or more processors for implementing the above-mentioned storage-computing integrated architecture simulator acceleration method based on dynamic domain decomposition.
[0044] An embodiment of the present invention also provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, it is used to implement the above-mentioned storage and computing integrated architecture simulator acceleration method based on dynamic domain decomposition.
[0045] An embodiment of the present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned storage-computing integrated architecture simulator acceleration method based on dynamic domain decomposition.
[0046] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0047] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0048] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the function specified in one process or more processes and / or boxes Figure 1 in one box or more boxes Figure 1 specified in the one process or more processes and / or boxes
[0049] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one process or more processes and / or boxes Figure 1 in one box or more boxes Figure 1 specified in the one process or more processes and / or boxes
[0050] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed in the present invention are within the protection scope of the present invention.
Claims
1. An accelerated method for a memory-computation integrated architecture simulator based on dynamic domain decomposition, characterized in that Including: Based on a compiler, convert the target network into a form suitable for execution on the in-memory computing architecture, that is, a binary instruction program for simulation, and additionally create a mapping file; the mapping file includes weight replication information, weight and in-memory computing device array mapping information, and the number of input times corresponding to each weight block; Obtain a configuration file, the simulator initializes and instantiates a simulation model based on the configuration file, and then performs domain decomposition on the simulation model based on the mapping file, where weight replication is used to balance the execution time between different network layers, and the number of input times corresponding to each weight block is used to keep the workload of each thread balanced, and each input of each weight block is equivalent to a unit workload of the simulator; The simulator assigns computing threads to each decomposed part, loads the binary instruction program, and the simulator starts the simulation task.
2. A method for accelerating a memory-computation integrated architecture simulator based on dynamic domain decomposition according to claim 1, characterized in that Also including: Measure the performance of the simulation process based on evaluation metrics, and the evaluation metrics include latency, energy consumption, power consumption, and accuracy.
3. A method for accelerating a memory - in - computing architecture simulator based on dynamic domain decomposition according to claim 1, characterized in that, The target network includes a deep neural network model.
4. A method for accelerating a memory - in - computing architecture simulator based on dynamic domain decomposition according to claim 1, characterized in that Based on a compiler, convert the target network into a form suitable for execution on the in-memory computing architecture, that is, a binary instruction program for simulation, and additionally create a mapping file, including: First, the compiler performs operator fusion; then the weight mapping process assigns the weights of the target network to the in-memory computing device array; subsequent data flow scheduling focuses on optimizing the flow path of data between different computing units; finally, the compiler generates a binary instruction program for simulation and additionally creates a mapping file.
5. A method for accelerating a memory - in - computing architecture simulator based on dynamic domain decomposition according to claim 1, characterized in that, The simulator assigns computing threads to each decomposed part, including: When assigning computing threads, combine the number of input times corresponding to each weight block to keep the workload of each thread balanced; or assign the computing cores mapping the same network layer to the same thread.
6. The accelerated method for a memory-computation integrated architecture simulator based on dynamic domain decomposition according to claim 5, characterized in that The computing cores assigned to the same thread share a time queue for simulation.
7. A method for accelerating a memory - in - computing architecture simulator based on dynamic domain decomposition according to claim 5, characterized in that, Assign the computing cores mapping the same network layer to the same thread, specifically: assign the n network layers to n threads respectively; to ensure the execution order of instructions, the simulation process is as follows: Thread 1 containing the computing core for the first layer weight calculation starts simulation from time 0; when reaching time a, the number of generated results reaches the threshold, and the event queue of thread 1 will send this data to thread 2 containing the computing core for the second layer weight calculation through inter-thread events; thread 2 will start simulation from time 0 only after receiving this data, and will synchronize with the data sent by thread 1 after its own simulation time reaches time a, and load this data into its own computing core; this ensures that thread 2 receives the data at the correct time; after that, thread 2 will pause at time a until thread 1 sends data again for simulation and synchronize at the moment of receiving new data; The interaction process between thread n - 1 and thread n is the same as the interaction process between thread 1 and thread 2.
8. An accelerator device for a memory-computation integrated architecture simulator based on dynamic domain decomposition, characterized in that Including one or more processors for implementing an in-memory computing architecture simulator acceleration method based on dynamic domain decomposition according to any one of claims 1 - 7.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When executed by a processor, the program is used to implement an acceleration method for a memory-computation integrated architecture simulator based on dynamic domain decomposition according to any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements an acceleration method for a memory-computation integrated architecture simulator based on dynamic domain decomposition according to any one of claims 1-7.
Citation Information
Patent Citations
Heterogeneous architecture parallel programming model optimization system
CN117032647A
Method and apparatus for leveraging simultaneous multithreading for bulk compute operations
US20230205692A1