Calculation and memory access total delay simulation method and related device
By dividing DRAM memory fetching into free loading/storage stages and dependency checking stages, and simulating DRAM FIFO behavior, the problem of insufficient efficiency and accuracy of existing simulation methods in complex application scenarios is solved, and more efficient and accurate calculation and memory fetching total delay simulation is achieved.
Patent Information
- Application Number
- CN202510102889.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
Existing total delay simulation methods for computing and memory access are insufficient in complex application scenarios, especially when frequent dependencies on checking and processing of a large number of different inputs are required.
A calculation and memory access total latency simulation method is proposed, which divides DRAM access into free load/storage stages and dependency check stages, and simulates DRAM FIFO behavior to record storage occupations, thereby quickly evaluating the calculation and memory access total latency.
By carefully simulating the memory access process, we capture data dependence and memory access conflicts, accurately evaluate the impact of memory access on DRAM bandwidth and storage resources, improve simulation efficiency and accuracy, and provide better computing and memory access solutions.
Smart Images

Figure CN120011192A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a simulation evaluation method, and specifically to a calculation and memory access total delay simulation method and related devices. Background Art
[0002] In the field of computer and information systems, accurate simulation of total computation and memory access delay is an important basis for evaluating system performance, optimizing algorithm design, and designing hardware architecture. However, existing simulation methods for total computation and memory access delay have exposed some significant problems in practical applications, which limit the efficiency and accuracy of simulation, especially in complex and changing application scenarios.
[0003] First, existing methods lack a free load / store phase. In the process of computing and memory access, data loading and storage are frequent operations, and these operations often have certain flexibility and non-strict timing dependencies. However, existing simulation methods usually fail to fully consider this flexibility and perform a large number of unnecessary dependency checks. These dependency checks not only increase the computational overhead of the simulation, but may also cause deviations between the simulation results and the behavior of the actual system.
[0004] Secondly, existing methods for simulating total computation and memory access delays are inefficient and may lead to a significant decrease in simulation speed, especially when dealing with a large number of different inputs or performing long simulations, as each input change may require the entire system to be re-simulated to evaluate its performance in scenarios where a large number of different inputs need to be simulated, such as searchers and optimization algorithms. Summary of the invention
[0005] The present application provides a method for simulating the total delay of computing and memory access and a related device to address the technical problems that the existing method for simulating the total delay of computing and memory access performs a large number of unnecessary dependency checks, as well as low simulation efficiency and simulation speed when the inputs are different.
[0006] In order to achieve the above objectives, this application adopts the following technical solutions: In a first aspect, the present application proposes a method for simulating total latency of computing and accessing memory, comprising: Determine the access time and order of each tensor on DRAM and the time of each computing slice based on the neural network to be run; Allowing DRAM access to be performed on each computing chip in sequence according to a free load / store phase and a dependency check phase; the dependency check phase includes checking data dependencies between load and store operations; Simulate DRAM FIFO behavior based on the free load / store phase and dependency check phase, and record the corresponding storage occupancy; Based on the recorded storage occupancy, the total latency of computation and memory access is obtained.
[0007] Furthermore, simulating DRAM FIFO behavior according to the free load / store phase and the dependency check phase also includes simulating DRAM FIFO behavior according to rules of tensor processing, DRAM operation and computing time balance.
[0008] Furthermore, the rules for balancing tensor processing, DRAM operations, and computing time include: (1) When the starting position of a tensor is less than or equal to the current processing position, the tensor can be processed; (2) If the tensor currently being processed is the output feature map of the current computing slice, the DRAM operation is performed after the output feature map is calculated; (3) If rules (1) and (2) are not satisfied and the DRAM time will exceed the computing time, the free load / store phase is stopped and the dependency check phase of the next computing slice is entered; (4) If the next computing slice has a dependency and the DRAM operation time exceeds the computing time after processing the dependency, the current calculation is paused.
[0009] Furthermore, the method further includes updating the changed part by using a coarse-grained local update method and / or a fine-grained local update method when performing at least two simulations.
[0010] Furthermore, the coarse-grained local update method includes: When the neural network to be run changes locally, the relative time at the local change is updated, and the total delay time after the local change is obtained by adding up the relative times of all local changes.
[0011] Furthermore, the method for fine-grained local update includes: When there is a change in the dependency relationship in the neural network to be run, the corresponding DRAM identifier and calculation identifier are recorded. During the next evaluation, when the same DRAM identifier and calculation identifier as the previous evaluation are encountered, the update is stopped, and the total delay time after the dependency relationship changes is obtained based on the changed calculation time.
[0012] Furthermore, the method for simulating DRAM FIFO behavior includes: Considering the computation time of the computation slice, space is reserved in the Buffer for all tensors whose starting position is equal to the current position pos; Goes through the free load / store phase and the dependency checking phase; Remove all tensors whose end position is equal to pos+1 from the Buffer.
[0013] In a second aspect, the present application proposes a total delay simulation system for computing and accessing memory, comprising: The basic condition determination module is used to determine the memory access time and memory access order of each tensor on DRAM and the time of each computing slice according to the neural network to be run; A phase division module, used to make DRAM access to each computing chip be executed in sequence according to a free load / store phase and a dependency check phase; the dependency check phase includes checking data dependencies between load and store operations; A simulation module, used to simulate DRAM FIFO behavior according to the free load / store phase and the dependency check phase, and record the corresponding storage occupancy; The computing module is used to draw the operation diagram of the neural network to be run according to the recorded storage occupancy, and obtain the total delay of computing and memory access.
[0014] In a third aspect, the present application proposes an electronic device, comprising: a memory, and one or more processors; the memory is coupled to the processor; wherein computer program code is stored in the memory, and the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the steps of the above-mentioned calculation and memory access total delay simulation method.
[0015] In a fourth aspect, the present application proposes a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned calculation and memory access total delay simulation method are implemented.
[0016] Compared with the prior art, this application has the following beneficial effects: The present application proposes a method for simulating the total delay of computing and memory access. When the memory access time and memory access order of each tensor and the time of each computing slice are known, the DRAM access is divided into two stages on each computing slice: the free load / store stage and the dependency check stage. Then, the DRAM FIFO behavior is simulated according to the free load / store stage and the dependency check stage, and the corresponding storage occupancy is recorded for rapid evaluation of the total delay of computing and memory access. By dividing the DRAM memory access into two stages, the actual memory access process can be simulated more carefully, which helps to capture complex behaviors in the memory access process, such as data dependencies and memory access conflicts. Simulating DRAM FIFO behavior can accurately evaluate the impact of the memory access process on DRAM bandwidth and storage resources. Through the simulation method of the present application, fast simulation can be achieved, and a better computing and memory access scheme can be obtained through simulation, which provides important performance indicators for optimizing the design and execution of neural networks and helps to understand the actual hardware design and algorithm optimization.
[0017] The present application also proposes a total delay simulation system for computing and memory access, an electronic device and a computer storage medium, which possess all the advantages of the total delay simulation method for computing and memory access. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 A schematic diagram of the first flow chart of the total memory access delay simulation method for calculating the present application; Figure 2 A schematic diagram of the principle of simulating DRAM FIFO behavior in an embodiment of the present application; Figure 3 This is a schematic diagram of a neural network layer to be run before optimization in an embodiment of the present application; Figure 4 This is a schematic diagram of a neural network layer to be run after optimization in an embodiment of the present application; Figure 5 Schematic diagram of three dependence modes in the embodiments of the present application; Figure 6 A schematic diagram of a DRAM and a Buffer operation in an embodiment of the present application; Figure 7 A schematic diagram of a system for simulating total memory access delay calculation for this application. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.
[0021] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for which protection is sought, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0022] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0023] In the description of the embodiments of the present application, it should be noted that if the terms "upper", "lower", "horizontal", "inner", etc. appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the drawings, or the orientation or position relationship in which the invented product is usually placed when used. It is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.
[0024] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", which does not mean that the structure must be completely horizontal, but can be slightly tilted.
[0025] In the description of the embodiments of the present application, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal connection of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0026] In the field of neural networks, with the continuous improvement of model complexity and the increasing diversification of application scenarios, higher and higher requirements are placed on the execution efficiency and performance evaluation of neural networks on hardware. Especially in the processing of large-scale neural networks, due to the massive data read and write operations involved, the memory access latency of DRAM (Dynamic Random Access Memory) becomes one of the key factors affecting the overall performance of neural networks. Traditional neural network performance evaluation methods often focus on simulation and optimization of computing, while relatively ignoring the impact of memory access on performance. However, in practical applications, memory access latency is often intertwined with computing latency, and together determines the execution efficiency of neural networks. Therefore, accurately simulating and evaluating the memory access behavior of neural networks is of great significance for optimizing neural network design and improving execution efficiency.
[0027] Therefore, there is an urgent need for an efficient neural network computing and memory access total delay simulation method that can accurately simulate the memory access process of the neural network, capture the complex behavior in the memory access process, and provide accurate computing and memory access total delay data.
[0028] Based on the above situation, the present application proposes a method for simulating total delay of calculation and memory access and related devices, and the present application is described in detail below in conjunction with embodiments and drawings.
[0029] like Figure 1 As shown, it is a first flow chart of the method for calculating and simulating the total memory access delay of the present application, which may include: S101, determining the memory access time and memory access order of each tensor on the DRAM, and the time of each computing slice according to the neural network to be run.
[0030] In practical applications, the neural network to be run can be comprehensively analyzed first, including its structure, such as the number of layers, the type of each layer, the size of the input and output tensors of each layer, etc. Then determine the storage location and size of each tensor in the neural network, such as weights, biases, inputs, outputs, etc., on DRAM. According to the execution process of the neural network, the memory access time of each tensor can be determined, that is, when a tensor needs to be read or written from DRAM. Then determine the memory access order, that is, which tensors need to be accessed first and which tensors can be accessed later according to the calculation order of the neural network. Subsequently, the calculation process of the neural network is divided into multiple calculation slices, each of which contains a part of the calculation task. The time required for each calculation slice is determined, which can usually be determined based on the complexity of the calculation task and the availability of calculation resources.
[0031] S102, DRAM access is performed on each computing chip in sequence according to a free load / store phase and a dependency check phase; the dependency check phase includes checking data dependencies between load and store operations.
[0032] In practical applications, at the beginning of each computing slice, the free load / store phase begins. In this phase, the DRAM controller can freely load (read) or store (write) tensors without being restricted by data dependencies, in order to maximize the use of DRAM bandwidth and improve data throughput. After the free load / store phase, the dependency check phase begins. In the dependency check phase, the DRAM controller needs to check the data dependencies between load and store operations to ensure the correctness and consistency of the data. For example, if a load operation depends on the result of a previous store operation, it is necessary to ensure that the store operation has been completed before the load operation can be performed.
[0033] S103, simulating DRAM FIFO behavior according to the free load / store phase and the dependency check phase, and recording corresponding storage occupancy.
[0034] It should be noted that FIFO (First In, First Out) is a common access method for DRAM. During the simulation process, DRAM access requests need to be processed according to the FIFO principle, that is, the first request to arrive is processed first. During the simulation process, the storage occupancy of DRAM also needs to be recorded in real time, which can include the storage location, size, and occupancy time of each tensor in DRAM, for subsequent calculation of the total delay.
[0035] S104, obtaining the total delay of calculation and memory access according to the recorded memory occupancy.
[0036] The recorded storage occupancy data is analyzed, including the access delay of each tensor, the bandwidth utilization of DRAM, etc. Based on the analyzed storage occupancy data and computing slice time, the total computing and memory access delay of the entire neural network can be calculated. Among them, the total delay usually includes computing delay (i.e., the time required to perform computing tasks) and memory access delay (i.e., the time required to access DRAM).
[0037] In practical applications, the overall latency can be reduced and the execution efficiency of neural networks can be improved by optimizing memory access order, reducing data dependencies, and improving DRAM bandwidth utilization.
[0038] The computing and memory access total delay simulation method of the present application divides the DRAM access into two stages for each computing slice, namely, the free load / store stage and the dependency check stage, given the memory access time and memory access order of each tensor on the DRAM, and the time of each computing slice, and then simulates the DRAM FIFO behavior and records the corresponding storage occupancy, so that the operation diagram of the entire network can be drawn, and the accuracy is very close to that of hardware simulation. It is of great significance for optimizing the execution efficiency of neural networks and improving system performance.
[0039] like Figure 2As shown, it is a schematic diagram of the principle of simulating DRAM FIFO behavior. In some embodiments of the present application, when simulating DRAM FIFO behavior, the calculation time of the calculation slice can be added first, and then space is reserved in the Buffer for all tensors whose starting positions are equal to the current position pos, and then the dynamic random access memory access (DRAM Access) of the free load / store phase and the dependency check phase is experienced, and finally all tensors whose end positions are equal to pos+1 are deleted from the Buffer. The calculation slice is a segment in the neural network calculation process, and each segment contains a part of the calculation task. The calculation time represents the time required to execute this calculation slice, which will affect the timing and order of subsequent DRAM access. Therefore, at the beginning of the simulation, the calculation time of the current calculation slice is considered first. Buffer is a temporary storage area for storing tensors that are about to be accessed or being processed. At each time point or position pos, the starting position of all tensors is checked. If the starting position of a tensor is equal to the current position pos, it means that this tensor is about to be accessed or processed, so it is necessary to reserve enough space for this tensor in the Buffer for subsequent loading or storage operations. The order of loading and storing tensors in DRAM is determined in the free load / store phase and the dependency check phase. Finally, at the end of each time point or position pos, the end position of all tensors is checked. If the end position of a tensor is equal to pos+1, it means that the tensor has finished being accessed or processed at the current time point and no longer needs to occupy space in the Buffer. The tensor needs to be deleted from the Buffer to free up space for other tensors. Through this process, the behavior of neural networks in the calculation and memory access process can be more accurately simulated, providing strong support for optimizing system performance.
[0040] In some embodiments of the present application, when simulating DRAM FIFO behavior, the following four rules may be followed: (1) When the starting position of a tensor is less than or equal to the current processing position, the tensor can be processed.
[0041] This rule determines when a tensor can be processed. During the simulation, each tensor has a starting position, which indicates when it becomes available or can be accessed. If the current processing position reaches or exceeds the starting position of the tensor, the tensor is considered "processable", that is, operations such as loading, calculation or storage can be performed.
[0042] (2) If the tensor currently being processed is the output feature map of the current computing slice, the DRAM operation is performed after the output feature map is calculated.
[0043] This rule applies to the output feature graph of a computing slice. The output feature graph is the result of the computing slice processing, and usually needs to be completely calculated before it can be stored or further processed. Therefore, when processing the output feature graph of a computing slice, you must wait for its calculation to be completed before performing subsequent DRAM operations, such as storing it in DRAM or using it as input for other computing slices.
[0044] (3) If rules (1) and (2) are not satisfied and the DRAM time will exceed the computation time, the free load / store phase is stopped and the dependency check phase of the next computation slice is entered.
[0045] This rule deals with the relationship between DRAM operations and computing time. In the free load / store phase, the DRAM controller can freely perform load and store operations. If there are currently no tensors that can be processed (that is, rule 1 is not satisfied) and the tensor currently being processed is not an output feature map (that is, rule 2 is not satisfied), and the DRAM operation time is about to exceed the computing time, then it is necessary to stop the free load / store phase and switch to the dependency check of the next computing slice. The dependency check is to ensure that all dependent tensors of the next computing slice are ready before the next computing slice starts executing.
[0046] (4) If the next computing slice has a dependency and the DRAM operation time exceeds the computing time after processing the dependency, the current calculation is paused.
[0047] This rule handles the dependencies between computing slices and the conflicts between DRAM operation time and computing time. If the next computing slice has dependencies, that is, it needs to wait for certain tensors or operations to complete before it can start executing, then after checking the dependencies and handling these dependencies, it is necessary to check whether the DRAM operation time has exceeded the computing time. If it exceeds, in order to maintain the stability and correctness of the system, it is necessary to pause the current calculation, wait for the DRAM operation to complete, or adjust the execution order of the computing slices.
[0048] These four rules together form the basic framework for simulating DRAM FIFO behavior, ensuring the orderly processing of tensors, the reasonable scheduling of DRAM operations, and the proper handling of dependencies between computing slices, thus providing strong support for simulating and optimizing the neural network computing process.
[0049] In practical applications, the neural network to be run can be optimized and adjusted based on the results of the calculation and total memory access delay simulation method of this application. After the optimization and adjustment, the simulation needs to be re-performed. In order to improve the simulation efficiency during the optimization process, when performing multiple similar simulations, the simulation method of this application can be further accelerated. Specifically, coarse-grained local updates and fine-grained local updates can be used according to the optimization situation: (1) Coarse-grained local update: This is achieved by recording relative time and directly replacing the changed part.
[0050] like Figure 3 As shown in, it is a schematic diagram of the neural network layer to be run before optimization, Figure 4 As shown, this is a schematic diagram of the neural network layer to be run after optimization. Figure 3 and Figure 4 The numbers in represent computation slices. When the change affects only a part, such as when the computational features of a layer in a neural network change, only the local relative time can be updated. Figure 3 and Figure 4 In the example, the granularity of a certain layer has changed, but it only affects all the computing slices and tensors related to this layer, and the rest does not need to be updated. In this case, only relative time can be recorded instead of absolute time. For example, this DRAM tensor only records the time from the start of the previous DRAM tensor transmission to the start of the current DRAM tensor transmission, rather than the absolute time when the current DRAM tensor transmission starts. Finally, if absolute time is needed, only the changed part can be replaced and all relative times can be added up.
[0051] The above coarse-grained local update method greatly reduces the time and computational cost of re-evaluating the performance of the entire model due to small changes. It also allows developers to experiment and adjust different parts of the model more flexibly, which is crucial for rapid iteration and optimization.
[0052] (2) Fine-grained local updates: achieved by checking the same effective dependencies.
[0053] like Figure 5 As shown, it is a schematic diagram of three dependency modes. From left to right, the start of COMP depends on the end of the DRAM block, the start of DRAM depends on the end of the COMP block, and the start of DRAM depends on the start of the COMP block. In this embodiment, the effective dependency refers to the dependency that causes STALL (computation or memory access is idle). Record all effective dependencies and corresponding (DRAM_id, COMP_id) tuples. Among them, DRAM_id is the DRAM identifier, and COMP_id is the calculation identifier. The next time you evaluate, if only the order or time of some DRAM and COMP blocks has changed, you can start directly from the first change. When you encounter the same effective dependency as the last evaluate, you can stop updating, that is, if the same (DRAM_id, COMP_id) tuple is found in STALL (idle state), you can stop.
[0054] This application can be used as an accurate evaluation tool to evaluate various scheduling schemes under different hardware configurations in terms of energy cost and latency. The evaluation process follows a local-to-global approach, first evaluating each compute block and DRAM load / store request (DRAM tensor) separately, and then evaluating it as a whole.
[0055] In practice, for each computational block, the input feature maps (ifmaps) and weights have been pre-fetched into the global buffer (GBUF), while the output feature maps (ofmaps) are written back to the GBUF. The evaluation method of this application evaluates each interaction between the GBUF and the buffer, as well as the computational process, while taking into account dependencies to evaluate the overall performance and energy consumption. The corresponding energy consumption and computational time of the optimal solution searched are regarded as the energy consumption and computational time of the computational block. The energy consumption of each DRAM communication tensor is calculated by multiplying the amount of read and written data by their respective unit energy consumptions and then adding these products. The read and write time is calculated by dividing the amount of data by their respective bandwidths. The total energy consumption is calculated by adding the energy consumption of the above subcomponents, which is similar to existing classic works. The total computational time is based on the evaluation time of all computational blocks and DRAM tensors, using the following method. For each DRAM tensor, it can only start execution when the following three conditions are met: 1) The previous DRAM tensor has been completed; 2) For input feature maps (ifmaps) or weights, their start time must be less than or equal to the ID of the current computation block; 3) For the output feature map (ofmaps), it must wait for the computation block that generates it to complete before it can start.
[0056] like Figure 6 As shown in FIG. 1 , a schematic diagram of DRAM and Buffer operation is shown. C2 The start time is C1, but the previous DRAM tensor (W D ) is not completed until E1, so it can only start in the middle of E1. In addition, W B The start time is B. Although the previous DRAM tensor (I A2 ) has completed, but it still has to wait for A2 to complete before it can start. Each computation block can only start executing if the following conditions are met: 1) All required data (input feature maps, weights, etc.) are ready. A1, B1, and C1 cannot immediately follow their respective previous computation blocks because the required data is not yet ready when the previous computation block ends.
[0057] 2) All DRAM tensors with end times less than or equal to this computation block must have completed. For example, D1 cannot be followed by E1 because OE1 The end time is D1, and D1 must wait until O E1 It can only start after execution is completed.
[0058] like Figure 7 As shown, it is a schematic diagram of a total delay simulation system for calculating and accessing memory in the present application, which may include: The basic condition determination module is used to determine the memory access time and memory access order of each tensor on DRAM and the time of each computing slice according to the neural network to be run; A phase division module, used to make DRAM access to each computing chip be executed in sequence according to a free load / store phase and a dependency check phase; the dependency check phase includes checking data dependencies between load and store operations; A simulation module, used to simulate DRAM FIFO behavior according to the free load / store phase and the dependency check phase, and record the corresponding storage occupancy; The computing module is used to draw the operation diagram of the neural network to be run according to the recorded storage occupancy, and obtain the total delay of computing and memory access.
[0059] It should be noted that in the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are merely schematic. For example, the division of each module is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another device, or some features can be ignored or not executed. The module described as a separate component may or may not be physically separated. The component displayed as a module may be a physical unit or multiple physical units, that is, it may be located in one place, or it may be distributed in multiple different places. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0060] In addition, each module in each embodiment of the present invention may be integrated into a processing unit, each module may exist physically separately, or two or more modules may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0061] An embodiment of the present application also provides an electronic device, which may include one or more processors, a memory, and a communication interface.
[0062] The memory, the communication interface and the processor are coupled, for example, the memory, the communication interface and the processor may be coupled together via a bus.
[0063] The communication interface is used for data transmission with other devices. The memory stores computer program code. The computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the steps of the above-mentioned calculation and access total delay simulation method.
[0064] Wherein, the processor can be a processor or a controller, for example, a central processing unit (CPU), a general processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the present disclosure. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of DSP and microprocessors, and the like. The processor can be used to support electronic devices to execute the method steps provided in the above embodiments.
[0065] The bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The above bus may be divided into an address bus, a data bus, a control bus, etc.
[0066] An embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned calculation and total memory access delay simulation method are implemented.
[0067] The computer-readable storage medium involved in the present application includes random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the technical field.
[0068] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for simulating total delay of computing and memory access, characterized in that: include: Determine the access time and order of each tensor on DRAM and the time of each computing slice based on the neural network to be run; Allowing DRAM access to be performed on each computing chip in sequence according to a free load / store phase and a dependency check phase; the dependency check phase includes checking data dependencies between load and store operations; Simulate DRAM FIFO behavior based on the free load / store phase and dependency check phase, and record the corresponding storage occupancy; Based on the recorded storage occupancy, the total latency of computation and memory access is obtained.
2. The method for simulating total memory access delay according to claim 1, characterized in that: The simulating DRAM FIFO behavior according to the free load / store phase and the dependency check phase also includes simulating DRAM FIFO behavior according to the rules of tensor processing, DRAM operation and computing time balance.
3. The method for simulating total memory access delay according to claim 2, characterized in that: The rules for balancing tensor processing, DRAM operations, and computation time include: (1) When the starting position of a tensor is less than or equal to the current processing position, the tensor can be processed; (2) If the tensor currently being processed is the output feature map of the current computing slice, the DRAM operation is performed after the output feature map is calculated; (3) If rules (1) and (2) are not satisfied and the DRAM time will exceed the computing time, the free load / store phase is stopped and the dependency check phase of the next computing slice is entered; (4) If the next computing slice has a dependency and the DRAM operation time exceeds the computing time after processing the dependency, the current calculation is paused.
4. The method for simulating total memory access delay according to claim 1, characterized in that: The method further includes updating the changed part by using a coarse-grained local update method and / or a fine-grained local update method when performing at least two simulations.
5. The method for simulating total memory access delay according to claim 4, characterized in that: The coarse-grained local update method comprises: When the neural network to be run changes locally, the relative time at the local change is updated, and the total delay time after the local change is obtained by adding up the relative times of all local changes.
6. The method for simulating total memory access delay according to claim 5, characterized in that: The fine-grained local update method comprises: When there is a change in the dependency relationship in the neural network to be run, the corresponding DRAM identifier and calculation identifier are recorded. During the next evaluation, when the same DRAM identifier and calculation identifier as the previous evaluation are encountered, the update is stopped, and the total delay time after the dependency relationship changes is obtained based on the changed calculation time.
7. The method for simulating total memory access delay according to claim 1, characterized in that: The method for simulating DRAM FIFO behavior comprises: Considering the computation time of the computation slice, space is reserved in the Buffer for all tensors whose starting position is equal to the current position pos; Goes through the free load / store phase and the dependency checking phase; Remove all tensors whose end position is equal to pos+1 from the Buffer.
8. A computing and memory access total delay simulation system, characterized in that: include: The basic condition determination module is used to determine the memory access time and memory access order of each tensor on DRAM and the time of each computing slice according to the neural network to be run; A phase division module, used to make DRAM access to each computing chip be executed in sequence according to a free load / store phase and a dependency check phase; the dependency check phase includes checking data dependencies between load and store operations; A simulation module, used to simulate DRAM FIFO behavior according to the free load / store phase and the dependency check phase, and record the corresponding storage occupancy; The computing module is used to draw the operation diagram of the neural network to be run according to the recorded storage occupancy, and obtain the total delay of computing and memory access.
9. An electronic device, characterized in that: include: A memory, one or more processors; the memory is coupled to the processor; wherein the memory stores computer program code, the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the steps of the method for simulating the total delay of calculating and accessing memory as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for simulating the total delay of calculating and accessing memory as described in any one of claims 1 to 7 are implemented.