Simulation device, simulation system and its simulation method, storage medium

The simulation device with an agent, management, and interconnection modules addresses the lack of efficient simulation frameworks for AI hardware accelerators, enhancing neural network processor design and verification.

JP2025521149AActive Publication Date: 2025-07-08BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024570637
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-31
Filing Date
2023-05-30
Publication Date
2025-07-08
Estimated Expiration
2043-05-30

AI Technical Summary

Technical Problem

Current AI hardware accelerators lack a simple and efficient simulation framework for architecture design and function verification, with existing simulators being complex, not designed for AI accelerators, and having high overhead.

Method used

A simulation device is provided that includes an agent module, management module, and interconnection module to simulate a neural network processor, facilitating efficient system simulation and verification.

Benefits of technology

The simulation device accelerates the design process of neural network processors by providing a simple and efficient simulation framework, reducing development restrictions and improving computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025521149000001_ABST
    Figure 2025521149000001_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a simulation device, a simulation system and its simulation method, and a non-transitory computer-readable storage medium. The simulation device is used to simulate a neural network processor, communicates with an object model that simulates the neural network processor, and is configured to be an agent of the object model in the simulation device, including an agent module, a management module set to manage the simulation device, and an interconnection module configured to communicatively connect the agent module and the management module. The simulation device receives a task sent from a host, sends task-related work information to the object model by the agent module, receives feedback information returned after the object model processes the work information by the agent module, and provides the feedback information to the host.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the priority of Chinese Patent Application No. 202210613143.4 filed on May 31, 2022, and hereby incorporates by reference the content disclosed in the above Chinese patent application as part of this application.

[0002] Embodiments of the present disclosure relate to a simulation device, a simulation system and its simulation method, and a non-transitory computer-readable storage medium.

Background Art

[0003] With the development of Artificial Intelligence, the amount of parameters of algorithm models has increased rapidly, and the demand for hash rate has been increasing. For conventional hardware architectures (e.g., CPU (Central Processing Unit) / GPU (Graphics Processing Unit)), due to considering the balance between different business needs at the architecture design stage, the hash rate provided for AI applications is limited. As a result, Domain-Specific Accelerator (DSA) emerged. The core idea of DSA is to do specialized things using dedicated hardware as well. Since DSA satisfies applications within one domain rather than one fixed application, DSA can meet the trade-off between flexibility and specialization.

Summary of the Invention

[0004] To introduce the concept in a simplified form, this summary part is provided. These concepts will be described in detail in the detailed description part for implementing the following invention. This summary part is not intended to identify the key features or essential features of the claimed technical solution, nor is it used to limit the scope of the claimed technical solution.

[0005] At least one embodiment of the present disclosure is a simulation device for simulating a neural network processor, including an object model that simulates the neural network processor, an agent module configured to be an agent of the object model in the simulation device, a management module configured to manage the simulation device, and an interconnection module configured to communicatively connect the agent module and the management module. The simulation device receives a task transmitted from a host, transmits work information related to the task to the object model by the agent module, receives feedback information returned after the object model processes the work information by the agent module, and provides the feedback information to the host.

[0006] At least one embodiment of the present disclosure is a simulation system including the simulation device according to any one of the embodiments of the present disclosure, an object model, and a host. The host is configured to obtain the task and transmit the task to the simulation device, and the object model is configured to process the work information to obtain the feedback information.

[0007] At least one embodiment of the present disclosure is a simulation method applied to the simulation system according to any one of the embodiments of the present disclosure, including steps of transmitting a task by the host, receiving and analyzing the task by the simulation device and transmitting work information related to the task to the object model, processing the work information by the object model to obtain the feedback information, and providing the feedback information to the host by the simulation device.

[0008] At least one embodiment of the present disclosure further provides a non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the simulation method described in any one of the above embodiments.

Brief Description of the Drawings

[0009] The above and other features, advantages, and aspects of each embodiment of the present disclosure will become more apparent by referring to the following detailed description of the invention in conjunction with the drawings. In all the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the elements and components are not necessarily drawn to scale.

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Detailed Description of the Invention

[0010] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the drawings. Although several embodiments of the present disclosure are shown in the drawings, the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, it should be understood that providing these embodiments is for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are merely illustrative and do not limit the protection scope of the present disclosure.

[0011] It should be understood that each step described in the method embodiments of the present disclosure may be executed in a different order and / or executed in parallel. Also, the method embodiments may include additional steps and / or may be executed by omitting the steps shown. The scope of the present disclosure is not limited in this regard.

[0012] As used herein, the term "including" and its variations are non-limiting "including", that is, "including... but not limited thereto". The term "based on" means "at least partially based on". The term "an embodiment" represents "at least one embodiment". The term "another embodiment" represents "at least one another embodiment". The term "several embodiments" represents "at least several embodiments". Related definitions of other terms are given in the following description.

[0013] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only for distinguishing different devices, modules or units, and are not for limiting the order or interdependence of the functions executed by these devices, modules or units.

[0014] It should be noted that the modifications of "one" and "a plurality" mentioned in the present disclosure are illustrative and not restrictive. As understood by those skilled in the art, unless otherwise clearly specified in the context, it should be understood as "one or a plurality".

[0015] The names of messages or information that interact between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0016] Through recent developments, the architecture design of AI hardware accelerators has become increasingly complex. Starting from the initial single-core shared memory, it has gradually evolved to the current homogeneous multi-core distributed memory. The core design in AI hardware accelerators has also evolved from a simple Very Long Instruction Word (VLIW) structure to a situation where two implementation architectures of instruction-driven data access, namely a Turing-complete Instruction Set Architecture (ISA) and Master-Slave (data block) and Data-Streaming (data stream), coexist. For the early architecture design exploration and later verification of AI hardware accelerators, a complete simulation system for simulating AI hardware accelerators is essential. That is, current AI hardware accelerators lack a simple and efficient simulation framework for architecture design and function verification.

[0017] The optimization policies related to the AI compiler mainly include operator fusion, splitting, IO / computation parallel scheduling, etc. These optimization policies constitute a huge optimization search space. The final execution efficiency of each optimization policy on the underlying accelerator is different, and a cost model is required to evaluate each optimization policy and reduce the optimization search space. The cost model needs to be accurate enough without introducing excessive overhead during the compilation stage. Currently, the open-source full-system simulation platform includes the gem5 simulator. The drawbacks of the gem5 simulator mainly include that the system is huge and complex, not designed to model AI accelerators, and has low execution efficiency. Also, the gem5 simulator has too large an overhead as a cost model. Therefore, currently, a high-performance cost model is required to evaluate the AI compiler.

[0018] At least one embodiment of the present disclosure provides a simulation device. The simulation device is used to simulate a neural network processor and includes an agent module, a management module, and an interconnection module. The agent module communicates with an object model that simulates the neural network processor and is configured to be an agent of the object model in the simulation device. The management module is configured to manage the simulation device, and the interconnection module is configured to communicatively connect the agent module and the management module. The simulation device receives a task sent from a host, sends work information related to the task to the object model by the agent module, receives feedback information returned after the object model processes the work information by the agent module, and provides the feedback information to the host.

[0019] The simulation device provided by an embodiment of the present disclosure can effectively implement the architecture design, exploration, and functional verification of a neural network processor by modeling the entire system simulation platform of the neural network processor, and can accelerate the structural and mechanical design process of the neural network processor, thereby avoiding the development of the neural network processor being restricted by hardware. In addition, this simulation device has a simple structure and is easy to implement.

[0020] At least one embodiment of the present disclosure further provides a simulation system, its simulation method, and a non-transitory computer-readable storage medium.

[0021] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings, but the present disclosure is not limited to these specific embodiments. To keep the following description of the embodiments of the present disclosure clear and concise, the present disclosure omits the detailed description of some known functions and known components.

[0022] FIG. 1A is a schematic diagram of the hardware architecture of a simulation device provided by at least one embodiment of the present disclosure, FIG. 1B is a schematic diagram of the hardware architecture of another simulation device provided by at least one embodiment of the present disclosure, and FIG. 2 is a schematic diagram of the hardware architecture of a simulation system provided by at least one embodiment of the present disclosure.

[0023] The simulation device provided by an embodiment of the present disclosure can be used to simulate a neural network processor, and the neural network processor can be realized in a hardware form, for example, it can be realized as a chip based on a programming language. For example, the neural network processor can be used to implement convolution operations, matrix operations, etc.

[0024] As shown in FIGS. 1A and 1B, in some embodiments of the present disclosure, the simulation device 100 may include an agent module 110, a management module 120, and an interconnection module 130. As shown in FIG. 2, the agent module 110 communicates with an object model 300 that simulates a neural network processor, and is configured to be an agent of the object model in the simulation device 100. That is, the object model 300 is used to simulate the functions of the neural network processor. The management module 120 is configured to manage the simulation device 100, and the interconnection module 130 is configured to communicatively connect the agent module 110 and the management module 120. It should be noted that the simulation device 100 shown in FIG. 2 is the simulation device shown in FIG. 1A, but the simulation device 100 in the simulation system 1000 may be the simulation device shown in FIG. 1B.

[0025] For example, the management module 120 can manage the agent module 110 to control the agent module 110 to execute the corresponding functions.

[0026] For example, as shown in FIG. 2, the simulation device 100 receives a task sent from the host 200, sends work information related to the task to the object model by the agent module 110, receives feedback information returned after the object model processes the work information by the agent module 110, and provides the feedback information to the host 200.

[0027] For example, the simulation device 100 is realized by a virtual simulation platform. For example, the simulation device 100 can be realized as a QEMU (Quick EMUlator) virtual platform (virt platform) or the like.

[0028] For example, the management module 120 is implemented by a virtual central processing unit. For example, in some embodiments, the virtual central processing unit may be a RISC (Reduced Instruction Set Computer RISC)-V (V represents the fifth generation of RISC) core.

[0029] For example, the task may be any task that needs to be executed by a neural network processor such as target identification, matrix operation (e.g., matrix multiplication).

[0030] For example, the communication between the agent module 110 and the host 200 can be carried out in at least one mode including, for example, the socket mode, the shared memory mode, and / or the message queue mode. As shown in FIG. 2, in some embodiments, the agent module 110 and the host 200 communicate in the socket mode, the shared memory mode, and the message queue mode. The socket mode may be the Unix domain socket (UDS) mode. The shared memory mode indicates communicating in the mode of a shared memory file. For example, the shared memory file may be the / dev / shm / ivshmem file in the host 200. The ivshmem file can realize sharing the memory area created by the host 200 among different QEMU processes. The message queue (message Queue) mode MQ1 can transmit messages in the first in first out (FIFO) mode.

[0031] For example, the communication between the agent module 110 and the host 200 realizes the communication between large amounts of data in the shared memory mode, and realizes the communication between small amounts of data in the socket mode and / or the message queue mode.

[0032] For example, communication can also occur between the agent module 110 and the object model 300 in at least one way. For example, communication between the agent module 110 and the object model 300 can be in a message queue manner. As shown in FIG. 2, in some embodiments, the agent module 110 and the object model 300 communicate in a message queue manner. For example, the agent module 110 sends task-related work information to the object model 300 in the message queue manner MQ2, and the object model 300 sends feedback information to the agent module 110 in the message queue manner MQ3. Similarly, the message queue manner MQ2 and the message queue manner MQ3 can transmit messages in a FIFO manner.

[0033] For example, the object model 300 can be connected to the simulation device 100 in a pluggable manner, that is, the simulation device 100 can simulate different object models. For example, different object models can simulate different neural network processors.

[0034] For example, the object model 300 can accelerate the computing speed of the neural network processor, save computing time, and improve computing efficiency.

[0035] For example, as shown in FIG. 2, the host 200 can obtain a task from an external device and send the task to the simulation device 100 in a socket manner.

[0036] For example, information such as parameters of the neural network processor and inputs and / or outputs related to tasks is accessed by the host 200 and the simulation device 100 in a shared memory manner to perform operations such as reading or writing. For example, the inputs and / or outputs related to a task can be determined based on the type of the task. For example, in some embodiments, the task may be to identify a target object in an image and feedback an image with the target object marked to the host 200. In this case, the input image may be the input related to the task, and the image identified by the object model 300 and with the target object marked may be the output related to the task.

[0037] For example, the simulation device 100 can send (e.g., synchronously) feedback information to the host 200 in a message queue manner.

[0038] For example, in some embodiments, as shown in FIG. 1B, the simulation device 100 can further include an input / output module 140. The input / output module 140 can interact with the host 200. For example, it communicates with the host 200 in a socket manner to receive a task. For example, the input / output module 140 is configured to send the task to the agent module 110 or the management module 120. For example, the input / output module 140 can include a buffer or the like to store the task.

[0039] For example, as shown in FIGS. 1A and 1B, the address space of the agent module 110 can include an agent register space Re and a model space Mem1.

[0040] For example, the agent register space Re is used to define the registers of the neural network processor. All the registers of the neural network processor are memory mapped. Memory mapped means that the registers and memory of the device are uniformly addressed, that is, by positioning the registers of the device with a memory address, the registers of the neural network processor are defined within the memory space and can be read and written by being positioned in the manner of a memory address.

[0041] For example, the agent register space Re can communicate with the host 200 and the object model 300 in a message queue manner.

[0042] For example, the model space Mem1 is used to store the parameters of the neural network processor and the inputs and / or outputs related to tasks, etc.

[0043] For example, in some embodiments, the model space Mem1 is shared by the agent module 110 and the host 200. In this case, for example, the model space Mem1 communicates with the host 200 in a shared memory manner.

[0044] For example, in some embodiments, as shown in FIG. 2, the model space Mem1 may be mapped as the memory space Mem2 in the object model. As can be seen from this, the model space Mem1 and the memory space Mem2 are in the same address space, and the actual memory is mapped to the shared memory file in the host 200. That is, the model space Mem1 can actually be accessed by the simulation device 100, the host 200, and the object model 300. For example, when the host 200 writes information such as the parameters of the neural network processor and the input and / or output related to the task into the model space Mem1, the object model 300 can directly read the information such as the input and / or output related to the task written in the model space Mem1 and perform related processing.

[0045] For example, in some embodiments, as shown in FIG. 2, the address space of the agent module 110 may further include a setting instruction space Mem3. The setting instruction space Mem3 is used to store the setting instructions of the registers of the neural network processor and the control instructions related to the task.

[0046] For example, in some embodiments, the agent module 110, the management module 120, and the interconnection module 130 can be mounted on a high-speed serial computer expansion bus (PCIe, peripheral component interconnect express) system. The PCIe system can include several device types such as a root complex (RC), a bridge, a switch, and an endpoint. The root complex is an interface between the CPU and the PCle bus. The bridge provides an interface with other buses (e.g., PCI or PCI-x, or yet another PCle bus), and is sometimes called bridge forwarding. The switch provides expansion or aggregation capabilities, enabling more devices to be connected to the ports of the PCle. The switch can function as a packet router, identifying which path a given packet needs to take based on the address or other routing information, and is a bridge from PCIe to PCIe. The endpoint is located at the end of the topology of the PCIe bus system and is generally considered an initiator of bus operations (similar to the host in the PCI bus) or a completer (similar to the slave in the PCI bus).

[0047] For example, in some embodiments, the interconnection module 130 may be a root complex in the PCIe system.

[0048] For example, in some embodiments, the agent module 100 may be an endpoint in a PCIe system. For example, the agent module 100 may include a plurality of base address register (BAR) spaces, and the plurality of base address register spaces may include a first base address register space and a second base address register space. In some embodiments, the plurality of base address register spaces may include BAR0 space to BAR5 space, the first base address register space may be the BAR0 space, and the second base address register space may be the BAR3 space.

[0049] For example, the agent register space Re corresponds to the first base address register space, for example, the BAR0 space, and the configuration instruction space Mem3 corresponds to the second base address register space, for example, the BAR3 space. The base address register space refers to the BAR space of PCIe. The simulation device is connected to the HOST via the PCIe interface as a device of the HOST. The HOST needs to map the space on the simulation device to the BAR space of PCIe in order to access the simulation device. If there are multiple independently accessible spaces in the simulation device, PCIe provides multiple BAR spaces for mapping.

[0050] For example, the configuration instruction space Mem3 is mapped to the shared memory file ivshmem shared by the agent module 110 and the host 200, and the shared memory file ivshmem is located in the host 200. That is, in the simulation device 100, all accesses to the BAR3 space are transferred to the shared memory file ivshmem on the host 200. The host 200 realizes the transmission of configuration instructions to the agent module 110 in the simulation device 100 by writing to the shared memory file ivshmem.

[0051] For example, in some embodiments of the present disclosure, as shown in FIG. 2, for the model space Mem1, the memory space Mem2, and the setting instruction space Mem3, their actual memories are mapped to a shared memory file in the host 200. Moreover, the model space Mem1 and the memory space Mem2 can be mapped to the same position in the shared memory file, and the position where the setting instruction space Mem3 is mapped to the shared memory file is different from the position where the model space Mem1 / memory space Mem2 is mapped to the shared memory file.

[0052] For example, the agent module 110 is configured to provide the contents in the agent register space Re and the model space Mem1 to the object model 300 based on the contents in the setting instruction space Mem3 to set and schedule the object model 300, and perform operation simulation. When the model space Mem1 can be mapped as the memory space Mem2 in the object model, the object model 300 can directly access the model space Mem1 to obtain necessary data such as data related to inputs of tasks.

[0053] For example, the host 200 can interrupt the agent module 110 in a socket manner, that is, the task executed by the object model 300 transmitted from the host 200 can be executed in an interrupt manner.

[0054] For example, in some embodiments, the agent module 110 may include an interrupt register, and the host writes notification information regarding a task to the interrupt register and notifies the management module 120 in an interrupt manner to execute the task. For example, the content of the process of the memory manager (runtime) in the host 200 communicating with the simulation device 100 in a socket manner is as follows. The memory manager performs a write operation on the interrupt register of the agent module 110 in the simulation device 100 so as to write notification information regarding a task. After the notification information regarding the task is written into the interrupt register of the agent module 110, the interrupt register interrupts the management module 120. After the agent module 110 interrupts the management module 120, the agent module 110 executes an interrupt processing program (which belongs to a part of the driver program of the agent module 110). Since the interrupt processing program is in the kernel mode (when a process falls into the kernel code and is executed by a system call, it is in the kernel execution mode (kernel code), and at this time, the privilege level is the highest), it is necessary to notify the scheduler in the user mode (when a process executes its own code, it is in the user execution mode (i.e., the user mode), and at this time, the privilege level is the lowest). For example, the scheduler in the user mode can be notified in a signal manner to receive a new task (i.e., the task sent from the host 200).

[0055] For example, as shown in FIGS. 1A and 1B, the simulation device 100 may further include a software module 150. The software module 150 includes an application program App and a virtual machine system (Guest OS). The application program App includes various tools Tools, library files Library, a scheduling device Scheduler, etc. The virtual machine system may include an operating system kernel and a driver.

[0056] For example, an operating system kernel is executed in the management module 120. For example, the operating system kernel may be Linux (registered trademark) 5.2 kernel or the like.

[0057] For example, the driver program of the agent module 110 is loaded into the operating system kernel and executed. For example, the driver program of the agent module 110 is loaded into the kernel in the form of a kernel module. A kernel module is a concept of an operating system, and a driver program is generally loaded by the operating system as a kernel module.

[0058] At least one embodiment of the present disclosure further provides a simulation system.

[0059] For example, as shown in FIG. 2, the simulation system 1000 may include a simulation device 100, a host 200, and an object model 300. For the communication method among the simulation device 100, the host 200, and the object model 300, reference may be made to the description in the embodiment of the above simulation device 100, and repeated descriptions will not be given for overlapping parts.

[0060] For example, the host 200 is configured to obtain a task and send the task to the simulation device 100. For example, the host 200 includes a memory manager, and the memory manager communicates with the simulation device 100 to transmit the task to the simulation device 100, and receives feedback information returned after the object model 300 processes the work information related to the task from the simulation device 100.

[0061] For example, the object model 300 is configured to process work information to obtain feedback information.

[0062] For example, in some embodiments, the object model 300 is used to simulate the functions of a hardware accelerator (e.g., a neural network processor) and can be modeled using the SystemC language. SystemC is a modeling platform composed of a series of C++ class libraries, to which a simulation kernel is added, and it can support hardware modeling at the system level, the behavioral description level, and the register transfer level.

[0063] Also for example, the object model 300 can also be modeled using the Verilog language.

[0064] For example, the abstraction level of the object model 300 is at least one of the algorithm level (ALM), the system architecture level (SAM), the transaction level (TLM), and the register transfer level (RTL).

[0065] For example, in some embodiments, the object model 300 can include execution units such as a matrix execution unit and a vector execution unit, storage-related units such as an on-chip static random-access memory (SRAM) and an SRAM controller, and a microcontroller unit (MCU), etc.

[0066] For example, in some embodiments, the object model 300 can model at least a portion of the read memory pipeline and the arithmetic pipeline of the neural network processor, and can statistically calculate the average execution time for each operation in the neural network processor. For example, taking the CPU as an example, the arithmetic pipeline represents an integer execution unit, a floating-point execution unit, etc. Correspondingly, the read memory pipeline represents an IO operation, and the IO operation represents a LOAD (load) / SAVE (store) execution unit, which is usually also called an LSU.

[0067] For example, the average execution time of each operation can represent the number of execution cycles of each stage (for example, the stage represents a pipeline stage), that is, the number of cycles to execute each operation. For example, by statistically calculating the average execution time for each operation, an approximate cycle that can be calculated can obtain an accurate model.

[0068] To improve the execution efficiency of the object model 300 as a cost-model, the object model 300 supports two modes: a functional mode / performance mode. In the performance mode, only the delay on the pipeline (for example, the pipeline represents a linear communication model of a pipeline segment for exchanging data) is simulated, and the actual operations corresponding to the read memory pipeline and the arithmetic pipeline are not executed. An accurate simulation cycle count can be obtained in the performance mode, and the execution speed is improved by one order of magnitude compared to the functional mode.

[0069] In an embodiment of the present disclosure, SystemC is used to model the delays in the read memory pipeline and the arithmetic pipeline of the object model 300, and the object model 300 meets the requirements of an AI compiler for a high-performance cost model. Thereby, the object model 300 can be used as the cost model of the AI compiler. As the cost-model of the AI compiler, the inference time performance index of the resent50 network on the chip can be significantly improved. For example, the inference time can be shortened.

[0070] At least one embodiment of the present disclosure further provides a simulation method applied to a simulation system, where the simulation system is the simulation system provided by any one of the embodiments of the present disclosure. For example, it is the simulation system 1000 shown in FIG. 2.

[0071] FIG. 3 is a schematic flowchart of a simulation method provided by at least one embodiment of the present disclosure.

[0072] As shown in FIG. 3, this simulation method may include the following steps S10 to S13.

[0073] In step S10, a task is sent by the host.

[0074] In step S11, the simulation device receives and analyzes the task, and sends the working information related to the task to the object model.

[0075] In step S12, the object model processes the working information to obtain feedback information.

[0076] In step S13, the simulation device provides the feedback information to the host.

[0077] For example, taking the simulation system 1000 shown in FIG. 2 as an example, step S10 is realized by the host 200. In step S10, the host 200 can receive a task from an external device and send the task to the simulation device 100. For example, in some embodiments, the task can be sent by the user to the host 200 via an external device.

[0078] For example, step S11 is realized by the simulation device 100. For example, the management module 120 in the simulation device 100 can analyze the task, and then the agent module 110 in the simulation device 100 sends the work information related to the task to the object model 300.

[0079] For example, step S12 is realized by the object model 300. In step S12, the object model 300 can process the work information to obtain feedback information, and the feedback information can be sent to the simulation device 100. For example, the feedback information can include the result after the object model 300 processes the work information. For example, when the task is the identification of a target in an image, the feedback information can include the probability of having the target in the image, etc.

[0080] For example, step S13 is realized by the simulation device 100. In step S13, the simulation device 100 receives the feedback information returned after the object model 300 processes the work information by the agent module 110.

[0081] Regarding the technical effects that the simulation method can achieve, reference can be made to the relevant descriptions in the embodiments of the above simulation device and simulation system, and the repeated parts will not be described repeatedly.

[0082] At least one embodiment of the present disclosure further provides a simulation device, which may include one or more memories and one or more processors. It should be noted that the components of the above simulation device are merely exemplary and not limiting. According to actual application needs, the simulation device may further have other components, and the embodiments of the present disclosure do not specifically limit this.

[0083] For example, one or more memories are used to non-temporarily store computer-executable instructions, and one or more processors are configured to execute the computer-executable instructions. When the computer-executable instructions are executed by one or more processors, one or more steps in the simulation method described in any one of the embodiments of the present disclosure are realized. For the specific implementation of each step of this simulation method and the related explanatory content, reference may be made to the embodiments of the above simulation method, and the overlapping parts will not be described in detail here.

[0084] For example, the processor and the memory can communicate directly or indirectly with each other.

[0085] For example, the processor and the memory can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. Mutual communication between the processor and the memory can also be realized via a system bus, and the present disclosure does not limit this.

[0086] For example, the processor and the memory can be provided on the server side (or cloud side).

[0087] For example, the processor can control other components in the simulation device to execute desired functions. The processor may be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc. The processor may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a tensor processing unit (TPU), or other programmable logic devices, discrete gates, or transistor logic devices, other forms of processing units having data processing capabilities and / or program execution capabilities, such as discrete hardware components. The central processing unit (CPU) may be, for example, of the X86 or ARM architecture.

[0088] For example, the memory may be a computer-readable medium and may include any combination of one or more computer program products. The computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or high-speed cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions can be stored in the computer-readable storage medium, and the processor can execute the computer-readable instructions to realize various functions of the simulation device. Various application programs, various data, etc. may be further stored in the storage medium.

[0089] Regarding the technical effects that the simulation device can achieve, reference can be made to the relevant descriptions in the embodiments of the above simulation method, and repeated descriptions will not be given for overlapping parts.

[0090] FIG. 4 is a schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as shown in FIG. 4, one or more computer-executable instructions 401 can be non-temporarily stored in the non-transitory computer-readable storage medium 40. For example, when the computer-executable instructions 401 are executed by a processor, one or more steps in the simulation method described in any one of the embodiments of the present disclosure can be executed.

[0091] For example, the non-transitory computer-readable storage medium 40 can be applied to the above-described simulation device. For example, the non-transitory computer-readable storage medium 40 can include the memory in the above-described simulation device.

[0092] For example, the description of the non-transitory computer-readable storage medium 40 can refer to the description of the memory in the embodiment of the simulation device, and the repeated parts will not be described repeatedly.

[0093] Hereinafter, referring to FIG. 5, FIG. 5 shows a schematic structural diagram of an electronic device 500 suitable for realizing the embodiments of the present disclosure. The electronic device 500 may be a terminal device (for example, a computer) or a processor, etc., and can be used to execute the simulation method of the above embodiments. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (abbreviated as PDA), tablet computers (abbreviated as PAD), portable media players (abbreviated as PMP), in-vehicle terminals (for example, in-vehicle navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, and smart home devices. The electronic device shown in FIG. 5 is only an example and does not limit the functions and usage ranges of the embodiments of the present disclosure.

[0094] As shown in FIG. 5, the electronic device 500 may include a processing device (for example, a central processing unit, an image processing device, etc.) 501, and can execute various appropriate operations and processes based on a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 further stores various programs and data required for operating the electronic device 500. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0095] Typically, devices such as an input device 506 including a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc., an output device 507 including a liquid crystal display (LCD), a speaker, a vibrator, etc., a storage device 508 including a magnetic tape, a hard disk, etc., and a communication device 509 can be connected to the I / O interface 505. The communication device 509 can enable the electronic device 500 to wirelessly or wiredly communicate with other devices to exchange data. Although FIG. 5 shows the electronic device 500 having various devices, it should be understood that it is not required to implement or include all of the shown devices. More or fewer devices may alternatively be implemented or included.

[0096] In particular, according to an embodiment of the present disclosure, the process described with reference to the above flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product including a computer program carried on a non-transitory computer-readable medium, the computer program including program code for executing the method shown in the flowchart so as to execute one or more steps in the simulation method described above. In such an embodiment, the computer program can be downloaded and installed from a network by the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the processing device 501 can be caused to execute the above functions limited in the simulation method of the embodiment of the present disclosure.

[0097] In the context of the present disclosure, a computer-readable medium may be a tangible medium that can include or store a program used in or coupled to an instruction execution system, apparatus, or device. The computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above, but is not limited thereto. More specific examples of the computer-readable storage medium can include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above, but is not limited thereto. In the present disclosure, the computer-readable storage medium may be any tangible medium that includes or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave carrying computer-readable program code. Such a propagated data signal can take various forms, including electromagnetic signals, optical signals, or any suitable combination of the above, but is not limited thereto. The computer-readable signal medium may be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can transmit, propagate, or transmit a program used by or in combination with an instruction execution system, apparatus, or device. The program code included in the computer-readable medium can be transmitted over any suitable medium, including wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above, but is not limited thereto.

[0098] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0099] Computer program code for performing the operations of the present disclosure can be compiled in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and further include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer by any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (e.g., connected via the Internet using an Internet service provider).

[0100] Flowcharts and block diagrams in the drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram can represent a module, program segment, or part of code that includes one or more executable instructions for realizing a given logical function. Note that in some alternative implementations, the functions assigned to the blocks may be realized in an order different from the order shown in the drawings. For example, two consecutively shown blocks may actually be executed basically in parallel, or depending on the function, may be executed in the reverse order. Note that each block in the block diagram and / or flowchart diagram, and combinations of blocks in the block diagram and / or flowchart, may be realized by a dedicated hardware-based system that executes a given function or operation, or may be realized by a combination of dedicated hardware and computer instructions.

[0101] The units described in the embodiments of the present disclosure may be realized in the form of software or in the form of hardware. For example, the name of a unit does not necessarily limit the unit itself in some cases.

[0102] The functions described above in this specification can be executed, at least in part, by one or more hardware logic components. For example, without limitation, typical types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic circuits (CPLDs), and the like.

[0103] In a first aspect, according to one or more embodiments of the present disclosure, a simulation device for simulating a neural network processor includes an object model that communicates with the neural network processor to be simulated, an agent module configured to be an agent of the object model in the simulation device, a management module configured to manage the simulation device, and an interconnection module configured to communicatively connect the agent module and the management module. Here, the simulation device receives a task transmitted from a host, transmits work information related to the task to the object model by the agent module, receives feedback information returned after the object model processes the work information by the agent module, and provides the feedback information to the host.

[0104] According to one or more embodiments of the present disclosure, communication between the agent module and the host is performed in a socket manner, a shared memory manner, and / or a message queue manner.

[0105] According to one or more embodiments of the present disclosure, the address space of the agent module includes an agent register space and a model space. The agent register space is used to define the registers of the neural network processor, and the registers of the neural network processor are positioned by memory addresses. The model space is used to store the parameters of the neural network processor and the inputs and / or outputs related to the task.

[0106] According to one or more embodiments of the present disclosure, the model space is mapped as a storage space in the object model.

[0107] According to one or more embodiments of the present disclosure, the model space is shared by the agent module and the host.

[0108] According to one or more embodiments of the present disclosure, the address space of the agent module further includes a configuration instruction space, and the configuration instruction space is used to store configuration instructions of registers of the neural network processor and control instructions related to the task.

[0109] According to one or more embodiments of the present disclosure, the agent module includes a plurality of base address register spaces, the plurality of base address register spaces include a first base address register space and a second base address register space, the agent register space corresponds to the first base address register space, and the configuration instruction space corresponds to the second base address register space.

[0110] According to one or more embodiments of the present disclosure, the configuration instruction space is mapped to a shared memory file shared by the agent module and the host, and the shared memory file is located in the host.

[0111] According to one or more embodiments of the present disclosure, the agent module is configured to provide the content in the agent register space and the model space to the object model based on the content in the configuration instruction space to configure, schedule, and perform operation simulation on the object model.

[0112] According to one or more embodiments of the present disclosure, communication between the agent module and the object model is performed in a message queue manner.

[0113] According to one or more embodiments of the present disclosure, the operating system kernel is executed in the management module, and the driver program of the agent module is loaded into the operating system kernel.

[0114] According to one or more embodiments of the present disclosure, the agent module includes an interrupt register, and the host writes notification information regarding the task to the interrupt register and notifies the management module in an interrupt manner to execute the task.

[0115] According to one or more embodiments of the present disclosure, the simulation device further includes an input / output module. Communication between the input / output module and the host is performed in a socket manner to receive the task, and the input / output module is configured to send the task to the agent module or the management module.

[0116] According to one or more embodiments of the present disclosure, the simulation device is implemented by a virtual simulation platform, and the management module is implemented by a virtual central processing unit.

[0117] In a second aspect, according to one or more embodiments of the present disclosure, there is provided a simulation system including the simulation device according to any one of the embodiments of the present disclosure, the object model, and the host. Here, the host is configured to obtain the task and send the task to the simulation device, and the object model is configured to process the work information to obtain the feedback information.

[0118] According to one or more embodiments of the present disclosure, the abstraction level of the object model is at the algorithm level, system structure level, transaction level, or register transfer level.

[0119] According to one or more embodiments of the present disclosure, the object model models at least a part of the read storage pipeline and the arithmetic pipeline of the neural network processor, and statistically calculates the average execution time for each operation in the neural network processor.

[0120] In a third aspect, according to one or more embodiments of the present disclosure, there is provided a simulation method applicable to the simulation system described in any one of the embodiments of the present disclosure, the method including: sending a task by the host; receiving and analyzing the task by the simulation device, and sending work information related to the task to the object model; processing the work information by the object model to obtain the feedback information; and providing the feedback information to the host by the simulation device.

[0121] In a fourth aspect, according to one or more embodiments of the present disclosure, computer-executable instructions are stored which, when executed by a processor, implement the simulation method described in any one of the embodiments of the present disclosure.

[0122] The above description is only an explanation of the preferred embodiments of the present disclosure and the technical principles used. As will be understood by those skilled in the art, the scope of the disclosure covered by the present disclosure is not limited to the technical solutions formed by specific combinations of the above technical features. Without departing from the concept of the above disclosure, other technical solutions formed by arbitrarily combining the above technical features or their equivalent features should also be covered simultaneously. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions disclosed in the present disclosure (but not limited thereto) should be covered.

[0123] Also, although the operations are described in a particular order, this should not be construed as requiring that these operations be performed in the particular order or sequence shown. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although the above discussion includes several specific implementation details, these should not be construed as limitations on the scope of the present disclosure. Some features described in the context of individual embodiments may be combined further to be implemented in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented in multiple embodiments individually or in any suitable sub-combination.

[0124] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Conversely, the specific features and acts described above are merely exemplary forms for implementing the claims.

[0125] The following points need to be further explained regarding the present disclosure.

[0126] (1) The drawings of the embodiments of the present disclosure relate only to the structures according to the embodiments of the present disclosure, and other structures can refer to normal designs.

[0127] (2) Unless conflicting, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0128] The above description is only a form for implementing the invention of the present disclosure, but the protection scope of the present disclosure is not limited thereto, and the protection scope of the present disclosure is based on the protection scope of the following claims.

Claims

1. A simulation device for simulating a neural network processor, comprising: an agent module configured to communicate with an object model that simulates the neural network processor and to act as an agent of the object model in the simulation device; a management module configured to manage the simulation device; an interconnection module configured to communicatively connect the agent module and the management module, wherein the simulation device receives a task transmitted from a host, transmits work information related to the task to the object model by the agent module, receives feedback information returned after the object model processes the work information by the agent module, and provides the feedback information to the host.

2. The simulation device according to claim 1, wherein communication between the agent module and the host is performed in a socket manner, a shared memory manner, and / or a message queue manner.

3. The address space of the agent module includes an agent register space and a model space, wherein the agent register space is used to define registers of the neural network processor, and the registers of the neural network processor are positioned by memory addresses, and the model space is used to store parameters of the neural network processor and inputs and / or outputs related to the task according to claim 1 or 2.

4. The simulation device according to claim 3, wherein the model space is mapped as a storage space in the object model.

5. The simulation device according to claim 3 or 4, wherein the model space is shared by the agent module and the host.

6. The address space of the agent module further includes a setting instruction space, wherein the setting instruction space is used to store setting instructions for registers of the neural network processor and control instructions related to the task according to any one of claims 3 to 5.

7. The agent module includes a plurality of base address register spaces, and the plurality of base address register spaces include a first base address register space and a second base address register space. The simulation apparatus according to claim 6, wherein the agent register space corresponds to the first base address register space, and the setting instruction space corresponds to the second base address register space. **Claim 8** The simulation apparatus according to claim 6, wherein the setting instruction space is mapped to a shared memory file shared by the agent module and the host, and the shared memory file is located in the host. **Claim 9** The simulation apparatus according to claim 6, wherein the agent module is configured to provide the contents in the agent register space and the model space to the object model based on the contents in the setting instruction space, set and schedule the object model, and perform operation simulation. **Claim 10** The simulation apparatus according to any one of claims 1 to 9, wherein communication between the agent module and the object model is performed in a message queue manner. **Claim 11** The simulation apparatus according to any one of claims 1 to 10, wherein an operating system kernel is executed in the management module, and a driver program of the agent module is loaded into the operating system kernel. **Claim 12** The simulation apparatus according to any one of claims 1 to 11, wherein the agent module includes an interrupt register, the host writes notification information related to the task to the interrupt register, and notifies the management module in an interrupt manner to execute the task. **Claim 13** Further includes an input / output module. Communication between the input / output module and the host is performed in a socket manner to receive the task. The simulation apparatus according to any one of claims 1 to 12, wherein the input / output module is configured to transmit the task to the agent module or the management module. **Claim 14** The simulation device is realized by a virtual simulation platform, and the management module is realized by a virtual central processing unit. The simulation device according to any one of claims 1 to 13.

15. A simulation system including the simulation device according to any one of claims 1 to 14, the object model, and the host, The host is configured to acquire the task and transmit the task to the simulation device, The object model is configured to process the work information to obtain the feedback information. The simulation system.

16. The simulation system according to claim 15, wherein the abstraction level of the object model is an algorithm level, a system structure level, a transaction level, or a register transfer level.

17. The object model models at least a part of the read storage pipeline and the operation pipeline of the neural network processor, and statistically calculates the average execution time for each operation in the neural network processor. The simulation system according to claim 15 or 16.

18. A simulation method applied to the simulation system according to any one of claims 15 to 17, A step of transmitting a task by the host, Receiving and analyzing the task by the simulation device, and transmitting work information related to the task to the object model, Processing the work information by the object model to obtain the feedback information, A step of providing the feedback information to the host by the simulation device, A simulation method including.

19. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the simulation method according to claim 18.

Citation Information

Patent Citations

  • Neural network processor verification method and device, electronic equipment and storage medium

    CN114118356A

  • Parallel Neural Processors for Artificial Intelligence

    JP2020533668A

  • Machine learning runtime library for neural network acceleration

    JP2020537784A