Verification method of neural network accelerator, simulation platform and storage medium
Through simulation model, the neural network accelerator chip is compiled and event queue generation is generated, and the number of execution cycles and energy consumption is determined, which solves the problem of low efficiency of traditional verification methods and achieves faster verification and optimization.
Patent Information
- Application Number
- CN202311653973.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-30
AI Technical Summary
When evaluating the performance, area and power consumption of neural network accelerator chips, the verification cycle is long and the efficiency is extremely low, making it difficult to achieve rapid iteration and optimization in the auxiliary design process.
The preset neural network is compiled through the simulation model, a running instruction image is generated, and the hardware module unit and performance simulation data activated by each event in the event queue in the simulation model is determined, thereby obtaining the running efficiency and operating power consumption of the simulation model.
It significantly reduces the verification cycle of neural network accelerator chip and improves verification efficiency, so that the simulation results of the simulation model can be obtained more quickly after the architecture design parameters are changed, meeting the needs of rapid iteration and optimization.
Smart Images

Figure CN120066910A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a verification method, a simulation platform, and a storage medium for a neural network accelerator. Background Art
[0002] With the breakthrough of artificial intelligence technology, neural network accelerators for deep learning have become a current technological hotspot. As the process size of CMOS integrated circuits shrinks, the complexity of the very large scale integrated circuits (VLSI) used in neural network accelerator chips increases exponentially, and the design cycle and verification cycle also increase accordingly. Generally, through EDA tools, the performance, area, and power consumption of the chip are analyzed, the RTL code of the chip is converted into transistor gate circuits, and based on the electrical characteristics of all transistor gate circuits within the verification cycle, relatively accurate evaluation information of the chip's performance, area, and power consumption can be given. Neural network accelerators provide the huge computing power required for the deep learning of neural networks, including iterative calculations with the same calculation type. In iterative calculations, the network layers of deep learning generally have dozens to hundreds of layers, and the number of identical iterations within a convolutional layer exceeds tens of thousands, and each iteration requires dozens to hundreds of cycles to complete. If traditional methods are used to evaluate the performance, area, and power consumption of neural network accelerator chips, the required verification cycle is very long and the efficiency is extremely low. Summary of the Invention
[0003] The main purpose of this application is to provide a verification method, a simulation platform, and a storage medium for a neural network accelerator, which are used to improve the verification efficiency of neural network accelerator chips and reduce the verification cycle of neural network accelerator chips.
[0004] In a first aspect, an embodiment of this application provides a verification method for a neural network accelerator, including:
[0005] Compiling a preset neural network according to a simulation model of the neural network accelerator to obtain a running instruction image of the simulation model;
[0006] Generating an event queue according to the running instruction image, and determining the execution cycle number and consumed energy of each event according to the hardware module units and performance simulation data activated by each event in the simulation model, where the performance simulation data is obtained by inputting different training events into the simulation model;
[0007] Obtaining the running efficiency and running power consumption of the simulation model according to the execution cycle number and consumed energy of each event.
[0008] Second aspect, an embodiment of the present application further provides a simulation platform, which includes a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing connection communication between the processor and the memory. When the computer program is executed by the processor, the steps of any verification method of the neural network accelerator provided by the embodiments of the present application are implemented.
[0009] Third aspect, an embodiment of the present application further provides a storage medium, which stores one or more programs that can be executed by one or more processors to implement the steps of any verification method of the neural network accelerator provided by the embodiments of the present application.
[0010] An embodiment of the present application provides a verification method for a neural network accelerator, including: compiling a preset neural network according to a simulation model of the neural network accelerator to obtain a running instruction image of the simulation model; generating an event queue according to the running instruction image, and determining the execution cycle number and consumed energy of an event according to the hardware module unit activated by each event in the event queue and the performance simulation data, where the performance simulation data is obtained by inputting different training events into the simulation model; obtaining the running efficiency and running power consumption of the simulation model according to the execution cycle number and consumed energy of each event. In the above process, based on the characteristic of the small number of calculation types of the neural network accelerator, different training events are input into the pre-modeled hardware module unit to obtain the performance simulation data of the hardware module unit. The preset neural network for verification is compiled into a running instruction image that matches the simulation model of the neural network accelerator, and then the running instruction is converted into an event queue. When each event in the event queue is executed by the simulation model, different from the previous execution process, it is not necessary to perform a complete operation of each event in the corresponding hardware module unit, but the execution cycle number and consumed energy of the event can be directly retrieved from the performance simulation data, so as to obtain the running efficiency and running power consumption of the simulation model. Therefore, after the architecture design parameters of the simulation model are changed, the simulation results of the simulation model can be obtained more quickly. Description of the Drawings
[0011] Figure 1 is a schematic block diagram of a hardware classification provided by an embodiment of the present application;
[0012] Figure 2 is a schematic block diagram of a simulation model of a neural network accelerator provided by an embodiment of the present application;
[0013] Figure 3 is a schematic flowchart of a method for building a simulation model provided by an embodiment of the present application;
[0014] Figure 4It is a schematic flowchart of a verification method for a neural network accelerator provided by an embodiment of the present application;
[0015] Figure 5 It is a schematic flowchart of an operation instruction reading method provided by an embodiment of the present application;
[0016] Figure 6 It is a schematic flowchart of an event reading method provided by an embodiment of the present application;
[0017] Figure 7 It is a schematic flowchart of a verification method for a neural network accelerator provided by an embodiment of the present application;
[0018] Figure 8 It is an input / output schematic diagram of a simulation platform provided by an embodiment of the present application;
[0019] Figure 9 It is a structural block diagram of a simulation platform provided by an embodiment of the present application. Detailed implementation manners
[0020] To enable those skilled in the art to better understand the technical solutions of the present application, the verification method for the neural network accelerator provided by the present application will be described in detail below with reference to the accompanying drawings.
[0021] Example embodiments will be described more fully hereinafter with reference to the accompanying drawings, but the example embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the application to those skilled in the art.
[0022] In the case of no conflict, the various embodiments of the present application and the features in the embodiments may be combined with each other.
[0023] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0024] The terms used herein are only for describing specific embodiments and are not intended to limit the present application. As used herein, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that when the present specification uses the terms "comprises" and / or "consists of", it specifies the presence of the stated features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their groups.
[0025] Unless otherwise defined, all terms (including technical and scientific terms) used herein shall have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries shall be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present application, and will not be interpreted as having an idealized or overly formal meaning unless expressly so defined herein.
[0026] The verification method for the neural network accelerator provided by the embodiments of the present application is used for the auxiliary design of the neural network accelerator to improve the verification speed of the performance, power consumption, and area of the neural network accelerator during the design process. Traditional chip design relies on commercial simulation software for accurate cycle-based simulation verification of the chip. Although this verification can most accurately give the verification results of the system's performance, power consumption, and area, when applied to neural network accelerators with high computing power, it cannot quickly obtain verification data after the architectural design parameters of the neural network accelerator are changed, resulting in slow verification speed and long verification cycle, and cannot achieve rapid iteration and optimization of the neural network accelerator during the auxiliary design process.
[0027] The calculations provided by the neural network accelerator are intensive calculations with the same computational characteristics. The circuits and algorithms used for calculation have similarity and batch characteristics. Therefore, based on the computational characteristics of the neural network, the verification process can be simplified on the basis of traditional simulation tools, thereby improving the verification efficiency of the neural network accelerator.
[0028] The verification method for the neural network accelerator provided by the embodiments of the present application is executed by a simulation platform. The simulation platform can be assumed to be on a server, where the server can be an independent server, a server cluster, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. Users can log in to the simulation platform through a terminal device connected to the server and draw the hardware architecture of the simulation model of the neural network accelerator on the simulation platform.
[0029] When designing the chip of the neural network accelerator, design requirements need to be considered. For example, computing power requirements and power consumption requirements determine the performance parameters of various hardware module units in the simulation model. Therefore, before the chip of the neural network accelerator, the hardware module units in the simulation model need to be parameterized to support design space exploration. The process of this parameterization is also the process of modeling the hardware module units.
[0030] Specifically, based on the computing characteristics of the hardware module units of the neural network accelerator, each hardware module unit of the neural network accelerator is abstracted and modeled. Please refer to Figure 1 , Figure 1 which shows a schematic block diagram of a hardware classification provided by an embodiment of the present application. As Figure 1 shown, the hardware classes include: object class, gemm class, alu class, sram class, dram class, and dma class. The object class is a base class defined for the commonalities of other hardware classes (gemm class, alu class, sram class, dram class, and dma class). Other hardware classes can inherit from the object class and extend attributes and override functions according to their own characteristics. The function descriptions of each hardware class are shown in Table 1.
[0031] Table 1 - Function description table of hardware classes.
[0032]
[0033] Before defining the object class, it is necessary to determine the commonalities of the gemm class, alu class, sram class, dram class, and dma class. In terms of connection relationships, there are some common attributes among the hardware classes, including: connection relationship, whether it is occupied. Thus, the base class members shown in Table 2 can be determined.
[0034] Table 2 - Base class member table.
[0035]
[0036]
[0037] Highly abstract the behaviors of the gemm class, alu class, sram class, dram class, and dma class. The common behaviors among the hardware classes include: read, write, and execute. Thus, the main member functions of the base class shown in Table 3 can be determined.
[0038] Table 3 - Main member function table of the base class.
[0039]
[0040] The base class members in Table 2 and the main member functions in Table 3 are all commonalities that can be inherited by the gemm class, alu class, sram class, dram class, and dma class. Based on these commonalities, the object class is defined.
[0041] Based on the inherited commonalities, classes such as gemm, alu, sram, dram, and dma also need to configure parameters for their own characteristics. For example, the sram class has configurable parameters such as bandwidth, size, and base address, and the gemm class has configurable parameters such as the size of the computing array, the bit width of the input data, the bit width of the output data, the bit width of the intermediate result data, and the data type. List the configurable parameters of each hardware class and the supported value ranges to obtain the configurable parameter table shown in Table 4.
[0042] Table 4 - Configurable Parameter Table of Hardware Types.
[0043]
[0044]
[0045] Each hardware class overloads and expands different base classes according to its own characteristics. For example, the gemm class expands the hardware type members. For example, the read behavior of the sram class is to retrieve the data at the corresponding position according to the read address and output it, the write behavior is to write the input data to the corresponding position according to the write address, and the execution behavior is empty; the read behavior of the GEMM class is to read and output the result of the matrix multiplication operation, its write behavior is to store the corresponding operands, and the execution behavior is divided into matrix multiplication calculation and decoding to obtain events such as reading operands, executing, and writing results according to whether the operands are valid. Similar implementations exist for the ALU class, ACCUM class, etc.
[0046] Exemplarily, please refer to Table 5, which shows a hardware type member table of the gemm class provided by an embodiment of the present application. As shown in Table 5, the matrices A, B, and C have the following meanings in the following formula:
[0047] C = A * B
[0048] Table 5 - Hardware Type Member Table of the gemm Class.
[0049] Hardware type member name of gemm class Member type Member description readcycles_ int Time (cycles) consumed for reading sram data writecycles_ int Time (cycles) consumed for writing sram data execycles_ int Time (cycles) consumed for executing calculations data_valid_ bool Whether the data is valid gemm_in_t Determined by inbytes_ Input data type gemm_out_t Determined by outbytes_ Output data type M_ int Number of rows of matrix A calculated / tilesize_ N_ int Number of columns of matrix B calculated / tilesize_ K_ int Number of columns of matrix A calculated / tilesize_ src_addr1_ uint64_t Address of the first element of matrix A oprand_A_ std::vector <byte> < / byte> Elements of matrix A src_addr2_ uint64_t Address of the first element of matrix B oprand_B_ std::vector <byte> < / byte> Elements of matrix B des_addr_ uint64_t Address of the first element of matrix C res_ std::vector <byte> < / byte> Elements of matrix C src_sram1_id_ object::KID ID of the sram where matrix A is located src_sram2_id_ object::KID ID of the sram where matrix B is located des_sram_id_ object::KID ID of the sram where matrix C is located
[0050] After the above process, each hardware module unit is highly parameterized during modeling. The pre-modeled hardware module unit can configure parameters such as the number of computing units, the storage capacity, and the data access bandwidth according to different design requirements. Each hardware class can be transformed into the corresponding hardware module unit in the simulation platform to construct a simulation model. The architecture design parameters of the simulation model include: the functional type of the hardware module unit, the data processing ability, and the data transmission relationship.
[0051] Please refer to Figure 2 , Figure 2 which shows a schematic block diagram of a simulation model of a neural network accelerator provided by an embodiment of the present application. As Figure 2As shown in the figure, the hardware architecture of the simulation model includes: a processing unit (RISCV), a control unit (decoding unit), a matrix multiplication operation unit (GEMM), a matrix multiplication operation cache unit (GEMM buffer), a vector processing unit (ALU), and a matrix accumulation unit (ACCUM), an input cache unit (Input buffer), an output cache unit (Outputbuffer), and a memory unit. The data transfer relationships between the various hardware module units are as shown by the arrows in Figure 2 the figure.
[0052] Based on the hardware architecture of the simulation model as shown in Figure 2 the figure, an architecture instruction set is edited. The architecture instruction set is used to define the instructions for computing and data transfer at the instruction level. Since neural network computing is batch iterative computing, the architecture instruction set generally includes the starting positions, lengths of computing and transferring data, as well as some variable configuration information, such as step size, mask, etc.
[0053] Please refer to Table 6, which shows a description table of an architecture instruction set provided by an embodiment of the present application. As shown in Table 6, the architecture instruction set for the simulation model includes: matrix operation instructions, vector operation instructions, data transfer instructions, and input / output instructions. Among them, the matrix operation instructions are used to implement matrix multiply-accumulate operations, the vector operation instructions are used to implement arithmetic operations and logical operations, the data transfer instructions are used to transfer data between SRAM and hardware module units, and the input / output instructions are used to transfer data and instructions between peripherals and the neural network accelerator.
[0054] Table 6 - Description table of the architecture instruction set.
[0055]
[0056] Through the above hardware framework and architecture instruction set, a simulation model of the neural network accelerator is built in the simulation platform. Based on the simulation model of the neural network accelerator, the performance of the neural network accelerator can be verified before the neural network accelerator is taped out to ensure that the neural network accelerator meets the design requirements.
[0057] Before verifying the neural network accelerator, it is also necessary to build a simulation model of the neural network accelerator. Please refer to Figure 3 , Figure 3 which shows a schematic flowchart of a method for building a simulation model provided by an embodiment of the present application. As shown in Figure 3 the figure, the specific steps of the method for building the simulation model include: S101 - S102.
[0058] S101. Obtain the architecture design parameters and architecture instruction set of the neural network accelerator.
[0059] Exemplarily, the user logs in to the server installed with the simulation platform through the terminal device, and inputs the architecture design parameters of the neural network accelerator into the terminal device. The terminal device transmits the architecture design parameters to the server. The server configures the pre-modeled hardware module units according to the architecture design parameters, and when configuring the hardware module units, calls the architecture instruction set of the hardware module units from the database to complete the construction of the simulation model.
[0060] S102. Build a simulation model of the neural network accelerator according to the pre-modeled hardware module units, architecture design parameters, and architecture instruction set.
[0061] Exemplarily, as Figure 2 shown, the hardware module units of the simulation model are the basic units of the neural network accelerator, and these hardware module units have been pre-modeled. The simulation platform calls the corresponding architecture instruction set in the database to endow these hardware module units with the performance of computing and data transfer, and configures parameters such as the number of computing units, storage capacity, and data access bandwidth of the hardware module units according to the architecture design parameters input by the user, and determines the data transmission relationship between each hardware module unit according to the architecture design parameters to obtain the simulation model of the neural network accelerator.
[0062] The above process is the preliminary preparation work for building the simulation model. Thus, the user can build a simulation model of the neural network accelerator on the simulation platform and verify the performance of the neural network accelerator through the simulation model.
[0063] Please refer to Figure 4 , Figure 4 which shows a schematic flow chart of a verification method for a neural network accelerator provided by an embodiment of the present application. As Figure 4 shown, the specific steps of the verification method for the neural network accelerator provided by the embodiment of the present application include: S201 - S203.
[0064] S201. Compile a preset neural network according to the simulation model of the neural network accelerator to obtain a running instruction image of the simulation model.
[0065] Exemplarily, during the process of simulating the neural network accelerator, a trained preset neural network is required as the running excitation of the neural network accelerator. The preset neural network can be imported by the user through the terminal device or stored in the database of the server. The preset neural network can be a single neural network or a combination of multiple neural networks of different scales.
[0066] Exemplarily, configure a neural network compiler according to the architecture design parameters, architecture instruction set, and hardware unit modules of the simulation model, and compile a preset neural network according to the configured neural network compiler to obtain the operation instruction set of the neural network accelerator. In the neural network compiler, the compilation realizes the conversion of the structure and parameters of the neural network from a high-level representation (Pytorch, Tensoflow, ONNX) to a micro-architecture instruction set, generating an operation instruction image; it should be noted that this part of the technology is an independent technical means and not within the scope of the application for protection, so no detailed elaboration will be made here.
[0067] S202. Generate an event queue according to the operation instruction image, and determine the execution cycle number and energy consumption of each event in the event queue based on the hardware module units activated by each event in the simulation model and the preset performance simulation data, where the performance simulation data is obtained by inputting different training events into the simulation model.
[0068] Exemplarily, input different training events into the simulation model to obtain performance simulation data. Specifically, the performance simulation data is obtained by testing each pre-modeled hardware module unit that has been modeled according to different configurations and different test events, and the test events can cover the events that may occur in the subsequent verification process. The performance simulation data records the execution cycle number and energy consumption required for each hardware module unit to complete each test event under different configurations.
[0069] Since the calculation characteristics of the neural network are few, and the circuits and algorithms used for calculation have similarity and batch characteristics, the event types in the training events are also predictable. When testing the hardware module units, only the data size of each type of hardware module unit needs to be modified, and according to the design requirements, the range of the data size of each type of hardware module unit can be deduced. Therefore, the performance simulation data can completely cover the events that may occur in the subsequent verification process.
[0070] Exemplarily, the simulation platform is provided with a decoder, which is used to interpret the instructions in the operation instruction image. After inputting the instructions in the operation instruction image into the decoder, the decoder sequentially obtains the instructions in the operation instruction image, sets the execution parameters and mutual dependencies of the instructions according to the characteristics of different instructions, generates events from these execution parameters and mutual dependencies, and writes the events into the event queue.
[0071] An event processor is also set in the simulation platform. An execution condition for determining whether an event can be executed is set in the event processor. The event processor polls the event queue to obtain target events that meet the execution conditions. Since the execution of a target event in the simulation model will activate the corresponding hardware module unit, the event processor is also used to determine the hardware module units activated by each target event in the simulation model. After the event processor determines the hardware module units activated by each target event, it matches the target event and the hardware module units activated by the target event in the performance simulation data to determine the number of execution cycles and the energy consumption required for the hardware module units to execute the target event. The number of execution cycles and the energy consumption are the number of execution cycles and the energy consumption required for the target event to execute in the simulation model.
[0072] A running instruction image is generated by the trained preset neural network to stimulate the neural network accelerator simulation model, simulating the training process of the preset neural network training, so as to obtain the corresponding performance data of the neural network accelerator during the training process.
[0073] S203. Obtain the running efficiency and running power consumption of the simulation model according to the number of execution cycles and the energy consumption of each event.
[0074] Exemplarily, the verification result of the simulation model is obtained by comprehensively calculating the execution cycle and energy consumption of each event in the event queue. The start time and end time of each event will be recorded, and the overlapping parts of different events will be excluded. The total number of cycles for all events to be executed serially is the total number of cycles. Usually, the running power consumption is the total energy consumption divided by the total number of execution cycles. The quotient obtained by dividing the number of execution cycles required for GEMM calculation by the total number of execution cycles of the simulation model is defined as the running efficiency of the neural network accelerator for running a certain neural network.
[0075] Through testing, for common neural networks, the verification result of the verification method of the neural network accelerator provided by the embodiments of the present application has an error less than 10% compared with the commercial software, but the simulation efficiency is increased by more than 5 times, meeting the requirements of rapid iterative development of accelerator design.
[0076] In the above process, based on the characteristic of the small number of calculation types of the neural network accelerator, different training events are input into the pre-modeled hardware module unit to obtain the performance simulation data of the hardware module unit. The preset neural network for verification is compiled into a running instruction image that matches the simulation model of the neural network accelerator, and then the running instruction is converted into an event queue. When each event in the event queue is executed by the simulation model, different from the previous execution process, it is not necessary to perform a complete operation of each event in the corresponding hardware module unit, but the execution cycle number and energy consumption of the event can be directly retrieved from the performance simulation data, so as to obtain the running efficiency and running power consumption of the simulation model. Therefore, after the architecture design parameters of the simulation model are changed, the simulation results of the simulation model can be obtained more quickly.
[0077] To introduce the technical solution of this application more clearly, the technical solution of this application will also be introduced through specific embodiments below. It should be noted that the specific embodiments are used to expand the description of the technical solution of this application, rather than limiting this application.
[0078] In some embodiments, the performance simulation data further includes area simulation data. Before obtaining the verification data of the neural network accelerator according to the execution cycle number and energy consumption of each event, the method further includes: determining the total area of the neural network accelerator according to a preset area function, architecture design parameters, and area simulation data.
[0079] Exemplarily, detailed and accurate area models of each hardware module unit are constructed by using Synopsys Design Compiler (DC) under different process libraries. By separately simulating the architecture design parameters of each hardware module unit, accurate area simulation data is obtained to determine the module area of each hardware module unit under different architecture design parameters. All possible and valuable architecture design parameters are traversed as much as possible to improve the scalability and configurability of the simulation model. In addition, this comprehensive process is completely automated, so it can be easily ported to different process nodes. For the architecture design parameters not considered, the embodiments of this application construct a parameter-area linear interpolation function, that is, a preset area function, for each functional component by using Matlab cftool, so that the area model can process the architecture design parameters of any hardware module unit. The area simulation data is obtained from these real comprehensive data.
[0080] After obtaining the area simulation data, the unit area of each hardware module unit can be determined from the area simulation data according to the preset area function and the architecture design parameters of each hardware module unit, so as to obtain the total area of the neural network accelerator. In this way, not only the total area of the neural network accelerator is determined, but also the unit area of each hardware module unit is determined, which is convenient for improving the layout rationality of each hardware module unit when laying out the neural network accelerator.
[0081] In some embodiments, determining the total area of the neural network accelerator according to the preset area function, architecture design parameters and area simulation data includes: determining the unit area of the hardware module unit according to the architecture design parameters and area simulation data; determining the adjustment coefficient of the preset area function according to the architecture design parameters; obtaining the total area of the neural network accelerator according to the preset area function, adjustment coefficient and unit area.
[0082] Exemplarily, when the neural network accelerator is constructed, that is, when each hardware module unit is initialized and registered, the unit area of each hardware module unit is obtained by indexing the area simulation data. If there is no corresponding area data for a specific architecture design parameter in the area simulation data, area evaluation is performed according to the parameter-area linear interpolation function and the specific architecture design parameter, and the specific architecture design parameter is used to calculate the adjustment coefficient of the preset area function. In addition, considering that there are a small number of components for connection between each hardware module unit in the real neural network accelerator, the preset area function also includes a prior polynomial function, and the total area after accumulation is fine-tuned through the prior polynomial function, and then the overall total area of the accelerator is finally given. The prior function provided by the embodiments of the present application is evaluated through the architecture design parameters of different hardware module units, which ensures the generalization ability of the area model.
[0083] In some embodiments, the event queue includes execution events and transmission events, each execution event is executed by a single hardware module unit, and each transmission event is executed by multiple hardware module units.
[0084] Exemplarily, the transmission event indicates that the event is of the data transmission type, and the hardware module units required to execute the transmission event are multiple, because the data is transmitted between multiple hardware module units. The hardware module unit required for an execution event is one, and one execution event only corresponds to one hardware module unit being activated.
[0085] By differentiating the types of events, it is beneficial to subsequently determine the hardware module unit for executing the event, and the execution cycle number and energy consumption of the event can be more accurately determined during the simulation process.
[0086] In some embodiments, the performance simulation data includes timing simulation data. Before determining the number of execution cycles and the consumed energy of an event based on the hardware module units activated in the simulation model by each event in the event queue and the performance simulation data, it further includes: inputting training events with different parameters into the hardware module units, where the event types in the training events include the types of each event in the event queue; obtaining the simulation cycles required for the hardware module units to execute the corresponding training events to obtain the timing simulation data.
[0087] Exemplarily, in order to establish accurate timing data for a neural network accelerator, first, accurate timing simulation needs to be performed on different hardware module units. The embodiments of the present application use Synopsys Verilog Compiler Simulator (VCS) to obtain the timing information of different hardware module units, construct all possible and representative micro-test events, covering computing and data transfer operations under various conditions. Then, each hardware module unit is simulated in VCS under different architecture design parameters and different micro-test events. By checking the waveform files generated by the simulation, the number of execution cycles required for each hardware module unit to execute different operations (i.e., micro-test events) is obtained, and thus the timing simulation data is constructed. Based on this real timing database, the embodiments of the present application use an event-driven simulation framework to obtain the performance evaluation of the neural network accelerator. As the events in the event queue are executed one by one, the required number of cycles can be accurately accumulated.
[0088] In some embodiments, the performance simulation data includes power consumption simulation data. Before determining the number of execution cycles and the consumed energy of an event based on the hardware module units activated in the simulation model by each event in the event queue and the performance simulation data, it further includes: inputting training events with different parameters into the hardware module units, where the event types in the training events include the types of each event in the event queue; determining the simulation circuit of the hardware module units in the simulation model according to the architecture design parameters; obtaining the switching activity information generated when the hardware module units execute the corresponding training events; and determining the power consumption simulation data according to the switching activity information and the simulation circuit.
[0089] Exemplarily, to accurately evaluate the power consumption of a neural network accelerator, a power consumption database of different hardware module units and accurate switching activity information are required. In the embodiments of the present application, test events with different parameters are input into hardware module units with different architecture design parameters to test the RTL-level simulation circuit of the hardware module units, so as to determine the switching power consumption, internal circuit power consumption, and leakage power consumption data values of the hardware module units, and obtain power consumption simulation data. The power consumption database of different hardware components includes: static power consumption (leakage power consumption) and dynamic power consumption (switching power consumption, internal circuit power consumption). For static power consumption, similar to area modeling, accurate static power consumption is obtained by separately synthesizing each hardware module unit. In addition, the dynamic power consumption of each hardware module unit when executing events with different parameters is obtained through the Primer Time tool.
[0090] This is achieved by first performing RTL-level simulation on the hardware module unit by inputting different event excitations to obtain the switching activity information file of the simulation circuit of the hardware module unit, and then inputting the switching activity information file and the source file of the simulation circuit into Primer Time for dynamic power consumption simulation.
[0091] In some embodiments, the process of determining the performance simulation data of the on-chip memory SRAM in the embodiments of the present application is as follows: 1. SRAM blocks with different bit widths and depths are designed using a Memory Compiler. 2. Considering that computing hardware modules such as GEMM and ALU require large data bit widths, small SRAMs are spliced into large-bit-width SRAMs, and complex data and address selection strategies are designed to ensure that data read and write operations are completed within one cycle. 3. Similar to the performance simulation data of other hardware module units, the present invention obtains the performance simulation data of the on-chip memory module SRAM for Memory Compilers with different processes.
[0092] In this way, possible and valuable architecture design parameters and operations are traversed as much as possible to improve the scalability and configurability of the power consumption model. These real power consumption data constitute the power consumption database of the embodiments of the present application.
[0093] In some embodiments, the running instructions in the instruction mirror are sorted according to the first order, and an event queue is generated according to the instruction mirror, including: sequentially reading the target running instructions according to the first order, where the target running instructions are the un-decoded running instructions in the instruction mirror; after detecting that the hardware module unit corresponding to the target running instruction is in an idle state, decoding the target running instruction into an event in the event queue.
[0094] Exemplarily, please refer to Figure 5 , Figure 5 which shows the schematic flow chart of a method for reading running instructions provided by the embodiments of the present application. As Figure 5As shown, the specific steps of the running instruction reading method include: S301 - S306.
[0095] S301. Read the running instruction.
[0096] Exemplarily, a decoder is set in the simulation platform. The decoder is used to decode the running instruction into an event in the event queue. The decoder reads the target running instruction in the first order. The target running instruction is the undecoded running instruction in the running instruction mirror.
[0097] S302. Determine whether the running instruction ends.
[0098] Exemplarily, the decoder determines whether the current target running instruction is the last unexecuted running instruction in the running instruction mirror. If not, execute S303. If so, execute S304.
[0099] S303. Instruction decoding.
[0100] Exemplarily, the decoder determines the hardware module unit corresponding to the target running instruction and checks whether the corresponding hardware module unit is in an idle state, that is, whether the tag (label) of the hardware module unit is occupied.
[0101] S304. Set the last event.
[0102] S305. Determine whether the event queue is full.
[0103] Exemplarily, if the event queue is full or a data conflict occurs, suspend the execution of S303 until there is an empty space in the event queue.
[0104] S306. Generate an event and write it into the event queue.
[0105] Exemplarily, if the decoder determines that the hardware module unit is in an idle state, it takes the target running instruction as an event according to the instruction nature and puts the events into an event queue sorted by execution time in sequence.
[0106] In some embodiments, the events in the event queue are sorted in the second order. Executing each event in the event queue in the simulation model corresponding to the neural network accelerator includes: reading the target event in the second order. The target event is the unexecuted event in the event queue; after detecting that the hardware module unit for executing the target event is in an idle state, calling the hardware module unit of the target event to execute the target event and setting the hardware module unit of the target event to an occupied state.
[0107] Exemplarily, please refer to Figure 6 , Figure 6 shows a schematic flowchart of an event reading method provided by an embodiment of the present application. AsFigure 6 As shown, the specific steps of the running instruction reading method include: S401 - S408.
[0108] S401. Poll the event queue.
[0109] Exemplarily, the simulation platform is provided with a simulator that always executes the first event in the priority queue, simulating the execution of the event with the nearest future time in reality. The simulation platform can also poll the events in the event queue according to the second order.
[0110] S402. Determine whether the event is executable.
[0111] Exemplarily, determine whether the event is executable according to whether the hardware module unit corresponding to the event is in an occupied state. The event that meets the execution condition will be executed (the first event can be executed by default). If it can be executed, a resource occupation flag will be set. Other events that need to depend on the occupied resources cannot be executed. If the resources required by other events are not occupied, they will be executed simultaneously.
[0112] S403. Set the occupation flag.
[0113] Exemplarily, change the tag label of the resource (the hardware module unit corresponding to the event) of the event to the occupied state.
[0114] S404. Execute the event.
[0115] Exemplarily, if the event is a data transmission event, call the read function of the first module to obtain a data block, and then call the write function of the second module to write the data block; if it is a hardware execution event, call the execution function of the first module, which can specifically perform operations or decode more events, determined by the execution function overloaded by each module and the validity of the data.
[0116] S405. Release the resources occupied by the event.
[0117] Exemplarily, after the event is executed, it will release the occupied resources, that is, change the tag label of the hardware module unit corresponding to the event to the idle state.
[0118] S406. Delete the event.
[0119] S407. Determine whether all events have been executed.
[0120] Exemplarily, after each event is executed, it will be judged whether the current time needs to be updated, and the next event will be executed in a loop until the event queue is empty.
[0121] S408. Output the verification result.
[0122] Exemplarily, after all events are executed, the overall performance and power consumption of the statistical system are counted, and a statistical report and the results of the final simulation calculation are output.
[0123] To more clearly illustrate the technical solution of the present application, the technical solution of the present application will also be elaborated through the following example, which is only used for explanation and is not used to limit the present application.
[0124] Please refer to Figure 7 , Figure 7 which shows a schematic flowchart of a verification method for a neural network accelerator provided by an embodiment of the present application. As Figure 7 shown, the specific steps of the verification method for the neural network accelerator include: S501 - S505.
[0125] S501. Implement hardware modeling through Chisel.
[0126] Exemplarily, each hardware module unit design is implemented using the Chisel hardware construction language. Since the Chisel language has relatively high abstraction and parameterization capabilities, it is convenient to quickly iterate and modify the hardware architecture of the neural network accelerator.
[0127] S502. RTL conversion.
[0128] Exemplarily, the hardware architecture of the neural network accelerator generated by Chisel is converted into RTL code through a compilation software.
[0129] S503. Commercial software simulation.
[0130] Exemplarily, a commercial AISC process tool is used for synthesis and RTL - level simulation to obtain information such as real - time timing, area, and power consumption under a specific process node.
[0131] S504. Mathematical fitting.
[0132] Exemplarily, for the on - chip cache SRAM, a corresponding RTL simulation file is obtained using a Memory Compiler. For DRAM, DRAMPower is used for simulation. The data obtained from fewer simulations is fitted using a Gaussian function, and the Gaussian function and corresponding parameters are written into the simulator, so as to obtain the timing, power consumption, and area metrics of modules with any configuration.
[0133] S505. Obtain verification data.
[0134] Exemplarily, please refer to Figure 8 , Figure 8 which shows an input - output schematic diagram of a simulation platform provided by an embodiment of the present application. As Figure 8As shown, the simulation platform includes: four inputs and one output. The four input data are respectively the hardware module modeling data, the performance simulation data, the hardware architecture of the neural network accelerator, and the running instruction image. One output is the verification result of the neural network accelerator, and this verification result includes the running efficiency, running power consumption, and area of the neural network accelerator. Among them, the running instruction image is obtained by compiling the hardware architecture of the neural network accelerator and a preset neural network through neural network compilation software.
[0135] The hardware architecture of the neural network accelerator is generated based on the hardware modeling data, architecture design parameters, and architecture instruction set. Thanks to the powerful parameterization configuration ability of Chisel, it can ensure that the configuration of hardware module units and corresponding architecture design parameters are scanned as much as possible to support the rapid design space exploration of the hardware architecture of the neural network accelerator. Through multiple loop scans, the running efficiency, running power consumption, and area under all architecture design parameters required for the neural network accelerator of deep learning can be obtained.
[0136] Please refer to Figure 9 , Figure 9 which shows the structural block diagram of a simulation platform provided by an embodiment of the present application. As Figure 9 shown, an embodiment of the present application proposes a simulation platform 30, which includes a processor 31, a memory 32, a computer program stored on the memory 32 and executable by the processor 31, and a data bus for realizing the connection and communication between the processor 31 and the memory 32. When the computer program is executed by the processor 31, it realizes the steps of the verification method of the neural network accelerator as provided in any one of the embodiments of the present application.
[0137] An embodiment of the present application provides a storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to realize the steps of the verification method of the neural network accelerator as provided in any one of the embodiments of the present application.
[0138] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the present application shall be within the scope of rights of the present application.
Claims
1. A verification method for a neural network accelerator, characterized in that, it includes: Compiling a preset neural network according to a simulation model of the neural network accelerator to obtain a running instruction image of the simulation model; Generating an event queue according to the running instruction image, and determining the execution cycle number and consumed energy of each event in the event queue according to the hardware module units activated in the simulation model by each event and the performance simulation data, where the performance simulation data is obtained by inputting different training events into the simulation model; Obtaining the running efficiency and running power consumption of the simulation model according to the execution cycle number and consumed energy of each event.
2. The verification method for a neural network accelerator according to claim 1, characterized in that, Before the step of compiling the preset neural network according to the simulation model of the neural network accelerator, the method further includes: Obtaining the architecture design parameters and architecture instruction set of the neural network accelerator; Building a simulation model of the neural network accelerator according to pre-modeled hardware module units, the architecture design parameters and the architecture instruction set.
3. The verification method for a neural network accelerator according to claim 2, characterized in that, The performance simulation data further includes area simulation data. Before obtaining the verification data of the neural network accelerator according to the execution cycle number and consumed energy of each event, the method further includes: Determining the unit area of the hardware module unit according to the architecture design parameters and the area simulation data; Determining the adjustment coefficient of the preset area function according to the architecture design parameters; Obtaining the total area of the neural network accelerator according to the preset area function, the adjustment coefficient and the unit area.
4. The verification method for a neural network accelerator according to claim 1, characterized in that, The event queue includes execution events and transmission events. Each execution event is executed by a single hardware module unit, and each transmission event is executed by multiple hardware module units.
5. The verification method for a neural network accelerator according to claim 1, characterized in that, The performance simulation data includes timing simulation data. Before determining the execution cycle number and consumed energy of each event according to the hardware module units activated in the simulation model by each event in the event queue and the performance simulation data, the method further includes: Inputting training events with different parameters into the hardware module units, where the event types in the training events include the types of each event in the event queue; Obtaining the simulation cycles required for the hardware module units to execute the corresponding training events to obtain the timing simulation data.
6. The verification method for a neural network accelerator according to claim 2, characterized in that, The performance simulation data includes power consumption simulation data. Before determining the execution cycle number and consumed energy of each event according to the hardware module units activated in the simulation model by each event in the event queue and the performance simulation data, the method further includes: Input training events with different parameters into the hardware module unit, where the event types in the training events include the types of each event in the event queue; Determine the simulation circuit of the hardware module unit according to the architecture design parameters; Obtain the switching activity information generated when the hardware module unit executes the corresponding training event; Determine the power consumption simulation data according to the switching activity information and the simulation circuit.
7. The verification method of the neural network accelerator according to claim 1, wherein, The operation instructions in the operation instruction mirror are sorted according to the first order, and generating the event queue according to the operation instruction mirror includes: Sequentially read the target operation instructions according to the first order, where the target operation instructions are the un-decoded operation instructions in the operation instruction mirror; After detecting that the hardware module unit corresponding to the target operation instruction is in an idle state, decode the target operation instruction into an event in the event queue.
8. The verification method of the neural network accelerator according to claim 1, wherein, The events in the event queue are sorted according to the second order, and executing each event in the event queue in the simulation model corresponding to the neural network accelerator includes: Sequentially read the target events according to the second order, where the target events are the un-executed events in the event queue; After detecting that the hardware module unit for executing the target event is in an idle state, call the hardware module unit of the target event to execute the target event, and set the hardware module unit of the target event to an occupied state.
9. A simulation platform, wherein, It includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing the connection and communication between the processor and the memory. When the program is executed by the processor, it implements the steps of the verification method of the neural network accelerator according to any one of claims 1-7.
10. A storage medium for computer-readable storage, wherein, The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the verification method of the neural network accelerator according to any one of claims 1-7.