Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

88 results about "Neural network hardware" patented technology

Pulse neural network hardware accelerator and data processing method

The invention discloses a pulse neural network hardware accelerator and a data processing method, and the accelerator is characterized in that a low-power-consumption three-stage pipeline CPU module is used for receiving input data, scheduling an SNN network acceleration instruction, and sending the input data to an asynchronous edge SNN hardware accelerator module through a coprocessor interface; the asynchronous edge SNN hardware accelerator module comprises a pulse data encoding and decoding module, L neuromorphic kernels and an on-chip network, the pulse data encoding and decoding module encodes input data into a pulse form and sends the pulse form into the neuromorphic kernels, and the neuromorphic kernels are used for performing calculation based on the data in the pulse form; the network-on-chip is used for communication between the neuromorphic kernels, and the connection between neurons before and after synapses in the neuromorphic kernels is realized by adopting a synaptic cross array. According to the invention, the data processing acceleration performance can be greatly improved.
Owner:WUHAN UNIV +1

Sparse binary neural network hardware accelerator for gesture recognition

The present invention belongs to the field of integrated circuit technology, and specifically is a sparse binary neural network hardware accelerator for gesture recognition. The accelerator of the present invention includes: an input cache module, a weight cache module, a data transmission on-chip network, 32 convolution calculation cores, a hierarchical accumulation tree, a subsequent processing unit, an HDMI interface, a UART interface, and a value prediction and sparse activation map compression / decompression module. The accelerator accelerates gesture recognition of sparse binary networks. The two-level value prediction method can skip the repeated calculations existing in the binary network, thereby accelerating the calculation cycle and reducing power consumption; the channel-level sparse activation compression and decompression method is used to reduce the required data storage volume; the present invention has the characteristics of low power consumption and low latency, effectively reducing hardware resources and improving the energy efficiency of the hardware.
Owner:FUDAN UNIVERSITY

Graph Neural Network Hardware Accelerator

The description relates to graph neural network hardware accelerators. One example can include multiple FPGAs or ASICs that each include multiple parallel arranged processing elements and a shared memory. Individual processing elements are configured to prune a subgraph of a graph neural network model. The shared memory is configured to recombine the pruned subgraphs to generate a pruned graph neural network model.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Dynamic convolution calculation method and architecture for hardware acceleration

The invention discloses a dynamic convolution calculation method and architecture for hardware acceleration, relates to the technical field of neural network hardware acceleration, and solves the technical problem that the current calculation method and architecture urgently need a hardware architecture capable of flexibly adjusting a convolution kernel, stride and filling parameters. The method comprises the steps of dynamically adjusting convolution calculation parameters in a convolution calculation task according to input data; decomposing the convolution calculation task into a plurality of subtasks according to the availability of hardware resources and the calculation amount of each convolution operation in the convolution calculation task; distributing the plurality of sub-tasks to different computing units according to a priority scheduling strategy; monitoring the load condition of each calculation unit in real time, and carrying out load balancing on each calculation unit; and executing each sub-task after load balancing in parallel through a plurality of calculation units to obtain a calculation result of the convolution calculation task. According to the invention, convolution calculation parameters can be flexibly adjusted, and the calculation efficiency and the resource utilization rate of the NPU are improved.
Owner:林培东 +2

Accelerating artificial neural networks using hardware-implemented lookup tables

The invention is notably directed to a hardware system (1) designed to implement an artificial neural network (ANN). The hardware system basically includes a neural processing apparatus (15), e.g., involving as crossbar array structure, one or more lookup table circuits (17), and one or more processing units (18). The neural processing apparatus is configured to implement M artificial neurons, where M≥1. The lookup table circuits are configured to implement a lookup table (LUT). The system further includes M′ processing units, where M≥M′≥1. Each processing unit is connected by at least one neuron, in order to be able to access a first value outputted by each connected neuron. In addition, each processing unit is connected to a LUT circuit, in order to efficiently access parameter values of a set of parameters from the LUT. Finally, each processing unit is configured to output a second value, corresponding to a value of a mathematical function taking said first value as argument. The mathematical function is otherwise determined by the set of parameters, the parameter values of which are accessed by each processing unit from the LUT, in operation. I.e., the mathematical function is defined (and thus determined) by a set of parameters, the values of which are efficiently retrieved from the hardware-implemented LUT. This results in a substantial acceleration of the computations of the function outputs, beyond the acceleration that may already be achieved within the neural processing apparatus and the processing units themselves. As a result, the neuron outputs can be more efficiently processed, prior to being passed to a next neuron layer. The invention is further directed to a method of operating such a hardware system.
Owner:AXELERA AI BV

Method for implementing in-memory computing with a small-signal ferroelectric capacitor based on non-destructive readout

The present invention proposes a method for realizing in-memory computing based on a small-signal ferroelectric capacitor with non-destructive readout, belonging to the field of novel computing architectures. The present invention realizes the linearly inseparable comparison operation of a content-addressable memory (CAM) unit and the local multiplication operation of a signed synaptic unit on a ferroelectric capacitor, reducing the hardware cost. Compared with the method of distance measurement based on the traditional CAM architecture, it has higher computational linearity and detection range, as well as lower search power consumption; at the same time, compared with the resistive synaptic array, it has no DC power consumption and no potential DC conduction path, avoiding the "IR drop" problem of large-scale arrays, and has lower computational power consumption, showing significant advantages in edge machine learning tasks based on feature retrieval and neural network hardware acceleration.
Owner:PEKING UNIV

ALU operation fusion processing module and method suitable for neural network

The invention discloses an ALU operation fusion processing module and method suitable for a neural network, and the module comprises a control unit which is used for receiving and decoding a machine instruction, managing the execution processes of internal and external circulation and microinstruction circulation, and generating a control signal of each stage of a microinstruction assembly line; the microinstruction buffer area is used for storing a microinstruction sequence pre-generated by the neural network compiler; the register file is used for storing source operands and results of ALU operation; each entry of the register file is composed of a valid bit, a tag bit and a data bit; the ALU computing core adopts an SIMD (Single Instruction Multiple Data) architecture and comprises a plurality of paths of parallel arithmetic logic function units; the Load / Store unit is used for processing data exchange between a register file and a local buffer area; and the data selection interface is used for selecting a data source or a target buffer area according to the storage tag field of the microinstruction. According to the method, the high efficiency and the flexibility of the ALU in the neural network hardware accelerator can be effectively considered.
Owner:ZHEJIANG UNIV

Convolutional neural network hardware acceleration method and system

The invention relates to the technical field of hardware acceleration, and provides a convolutional neural network hardware acceleration method and system, and the method comprises the steps: obtaining a convolutional neural network instruction which comprises an img2col instruction, a systolic array multiplication instruction and a col2img instruction; analyzing an operation instruction code and a function code field of each instruction, determining a target register, a source register and calculation parameters, determining an instruction function and generating a control signal; data to be operated are taken out from the source register and transmitted to the corresponding operation module together with the control signal and the calculation parameter to be calculated, and then the data are written back to the target register; through hardware optimization strategies such as multi-level address superposition, systolic array data multiplexing, double-buffer window construction and a parallel comparison tree, extra power consumption caused by multiple times of data loading and storage and intermediate result calculation is reduced, and data handling overhead, calculation delay and control complexity are greatly reduced.
Owner:SHANDONG UNIV

Softmax instruction set extension method and system based on RISC-V

The invention belongs to the field of neural network hardware acceleration, and provides a Softmax instruction set extension method and system based on RISC-V. The Softmax instruction set extension method based on the RISC-V. The Softmax instruction set extension method based on the RISC-V. The Softmax instruction set extension method based on the RISC-V. The Softmax instruction set extension method based on the RISC-V. The Softmax instruction set extension method comprises an instruction fetching stage, in the decoding stage, corresponding instruction functions are analyzed for Opcode, Funct7 and Funct3 of the Softmax instruction, and corresponding control signals are generated. In the execution stage, data to be subjected to Softmax operation is taken out from the data memory and transmitted into the Softmax calculation unit according to a control signal; executing Softmax calculation according to the following formula, and writing a Softmax calculation result back to the target register; the calculation process is divided into four sub-modules of maximum solution, index calculation, summation and normalization, and a hardware acceleration strategy is designed, so that the operation delay and the resource overhead are greatly reduced, the degree of parallelism, the precision and the energy efficiency ratio of calculation are effectively improved, and the method is suitable for a high-performance neural network reasoning acceleration scene.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Streamline processing graph to enhance performance of neural network execution on hardware accelerator

Neural networks have interconnected operations that process data, and the layered structure of neural networks enable complex machine learning tasks. Neural network hardware accelerators with specialized circuitry can perform the operations in an efficient manner. Some neural network hardware accelerators have a programmable look up table that can be used to efficiently apply an approximation of an activation function to an input tensor. To further improve efficiency, the programmable look up table can be used to store parameters that can be used to approximate a composite function that combines two or more activation functions. In some cases, the programmable look up table can be used to store two or more sets of parameters that can be used to approximate different two or more activation functions.
Owner:INTEL CORP +7

A hardware accelerator and acceleration method for convolutional neural networks

This invention discloses a convolutional neural network hardware accelerator, comprising: a memory that stores input feature map data in an input channel-first order; a feature map input module that reads input feature map data in an input channel-first reading order; a feature map caching module that caches the input feature map data in an input channel-first storage order; a convolution kernel module that stores convolution kernel data in an input channel-first storage order; a computation module that sequentially and in parallel reads M input feature map data and M convolution kernel data in each clock cycle, performs point-to-point multiplication on the input feature map data and the M convolution kernel data, accumulates all the point-to-point multiplication results to obtain the convolution result, and outputs the convolution result in parallel to N output channels as output feature map data; and a feature map output module that writes the output feature map data into the memory in an output channel-first storage order. This invention improves the computational efficiency of convolution operations in the hardware accelerator.
Owner:HANGZHOU FEISHU TECH CO LTD

A fully analog circuit system for handwritten digit recognition based on a storage-computing integrated architecture

The present invention discloses a fully analog circuit system for handwritten digit recognition based on a storage-computing integrated architecture, comprising a first analog circuit and a second analog circuit. The first analog circuit is composed of the following circuit modules: an input drive circuit, a first resistor array with a scale of m columns and 2n rows, a first current-voltage conversion circuit, a first subtraction circuit, and a first activation function circuit; the second analog circuit is composed of the following circuit modules: a second resistor array with a scale of n columns and 20 rows, a second current-voltage conversion circuit, a second subtraction circuit, a second activation function circuit, and a voltage comparison circuit. This system can eliminate the data handling problems brought about by the traditional von Neumann architecture and does not require the assistance of any digital modules, achieving ultra-low latency and accurate recognition. This system has very broad application prospects in the field of new neural network hardware system research, and is expected to provide a feasible non-von Neumann hardware solution for deep neural networks and edge computing.
Owner:ZHEJIANG UNIV

Storage and calculation integrated array based on self-rectification memristor and design method thereof

The invention relates to the technical field of neural network hardware implementation, and particularly discloses a storage and calculation integrated array of a self-rectification memristor and a design method of the storage and calculation integrated array. Comprising word lines, bit lines and self-rectification memristors, the word lines and the bit lines are perpendicular to each other, the self-rectification memristors are located at the intersections of the word lines and the bit lines, the positive ends of the self-rectification memristors are connected with the word lines, the negative ends of the self-rectification memristors are connected with the bit lines, the storage and calculation integrated array adopts a binary calculation mode, and non-0, namely 1 data and weights are used for calculation; the self-rectification memristors in adjacent rows or adjacent columns form a weight bit, if the weight is one bit weight, the calculation unit comprises one weight bit, and if the weight is n bit weight, the calculation unit comprises n weight bits; 0 or 1 is represented by the resistance states of the memristors, each bit weight is represented by the states of the two memristors at the same time, and the resistance states of the two self-rectification memristors at the same weight bit are opposite. According to the invention, calculation errors caused by non-linear resistance value reading during analog calculation of the self-rectification memristor can be eliminated.
Owner:SHANDONG UNIV

Neural network hardware device and system

Neural network systems and methods are provided. In one embodiment, a method of making a neural network device includes: forming a mesh layer on a substrate, the mesh layer including a matrix of randomly dispersed conductive nano-strands insulated from one another; forming an isolation trench extending into the mesh layer; forming a memristor device extending into the mesh layer, the memristor device including: an electrical conductor, and a layer of memristive material in electrical contact with individual nano-strands of a first set of conductive nano-strands in the mesh layer; forming an electrode extending into the mesh layer and spaced from the memristor device by the isolation trench, wherein the electrode is in electrical contact with individual conductive nano-strands of a second set of conductive nano-strands in the mesh layer; and forming a modulating device bridging the memristor device and the electrode.
Owner:THE GOVERNMENT OF THE UNITED STATES OF AMERICA AS REPRESENTED BY THE SECRETARY DEPARTMENT OF HEALTH & HUMAN SERVICES

Pipelined architecture based computing engine acceleration apparatus, method, device and medium

The application provides a pipeline structure-based computing engine acceleration device, method, equipment and medium in the technical field of neural network hardware acceleration, the device comprises: a circular row buffer, comprising: a first multiplexer, used for determining the storage position of n feature map pixels of an input feature map, wherein n is a positive integer greater than 0; m random access memories, used for caching n feature map pixels according to the storage position, wherein m is a positive integer greater than 0; a second multiplexer, used for outputting the feature map pixels written into the m random access memories; a stream driving filler, used for performing boundary filling on the feature map pixels output by the second multiplexer according to the position of the feature map pixels and the boundary parameter of the convolution layer in the neural network; a computing unit array, used for performing convolution calculation on the filled feature map pixels; and a weight memory, used for caching the weight data of the neural network.
Owner:INST OF SEMICONDUCTORS - CHINESE ACAD OF SCI

Hardware acceleration system and method supporting pulse-gated recurrent neural network

The invention discloses a hardware acceleration system and method supporting a pulse-gated recurrent neural network, and relates to the technical field of neural network hardware acceleration, and the system comprises a PC terminal and an accelerator. The accelerator comprises a UART (Universal Asynchronous Receiver / Transmitter) module, a data composer, a storage module, a controller module, a calculation module and a gating calculation core, the PC terminal is connected with the UART module; the gating calculation core is connected with the data composer; and the controller module and the storage module are respectively connected with the UART module, the data composer, the calculation module and the gating calculation core. Based on the above scheme, different coding modes can be selected at a PC terminal to adapt to different data types, and a pulse neuron activation function and a logic gate are utilized to perform calculation heterogeneous design in a calculation module and a gating calculation core, so that a high-performance configurable hardware acceleration calculation architecture for the pulse gating recurrent neural network is realized.
Owner:GUANGDONG UNIV OF TECH

Large language model softmax function hardware acceleration circuit and method

The invention discloses a large language model softmax function hardware acceleration circuit and method, and belongs to the field of neural network hardware acceleration of super-large scale integrated circuits. An input sequence is divided into a plurality of data blocks which are processed in parallel, and three-stage pipeline division is adopted, so that average single calculation delay is shortened to G clock cycles, the calculation parallelism is improved, the calculation speed of a softmax function is improved, and the reasoning delay is reduced; and a sparse mask strategy of sparse threshold comparison is introduced, the sparsity of data is fully utilized, the problems of high calculation complexity and high calculation delay of the softmax function and the data access bottleneck of the softmax function are solved, the calculation cost is remarkably reduced, and the calculation efficiency is improved. In addition, a softmax function hardware circuit adaptive to software optimization is constructed, calculation delay and memory access pressure are reduced through the hardware circuit, and data processing efficiency is improved.
Owner:NANJING UNIV

Low-power convolutional neural network hardware acceleration method based on stochastic computing

This invention discloses a low-power convolutional neural network hardware acceleration method based on random computation. First, a convolutional neural network model is constructed and trained. Then, during the inference phase, the trained model is deployed to a hardware computing unit, and random computation is used to partially or completely replace the convolutional computations in the model. The random computation includes three stages: encoding, computation, and decoding. In the encoding stage, the input feature map is scaled, and the convolutional kernel is quantized. The gray values ​​of the convolutional kernel and each point in the input feature map are normalized into probability values, and an encoding sequence is generated for each probability value. In the computation stage, the two encoding sequences of the convolutional kernel and corresponding points in the input feature map are subjected to a traversal AND operation to generate an AND operation sequence. In the decoding stage, the AND operation sequence is decoded to obtain the computation result. Without significantly affecting model performance, this method reduces computational complexity and hardware resource consumption, making it suitable for scenarios with limited power consumption and resources.
Owner:HEBEI UNIV OF TECH

Neural network hardware accelerator circuit with requantization circuits

A convolutional neural network includes convolution circuitry. The convolution circuitry performs convolution operations on input tensor values. The convolutional neural network includes requantization circuitry that requantizes convolution values output from the convolution circuitry.
Owner:STMICROELECTRONICS INT NV

A transconductance variable field effect transistor array and applications

The application discloses a transconductance variable field effect transistor array suitable for a dendritic network hardware and application, and belongs to the technical field of semiconductor integrated circuits.The application realizes three-element multiplication of a storage variable and two input variables based on a single transconductance variable field effect transistor, and realizes mapping of a dendritic network core algorithm based on a complementary device array.Compared with a traditional neural network hardware which realizes nonlinear transformation by using a neuron activation circuit, the application realizes nonlinear transformation by using intrinsic nonlinearity of a device, effectively reduces design complexity, optimizes area and power consumption of a system peripheral circuit, and has important significance for design of a high-performance artificial intelligence computing system.
Owner:PEKING UNIV

Neural network hardware implementation method and system based on incompletely specified function

The present invention provides a method and system for implementing neural network hardware based on an incompletely specified function. The method comprises: quantizing the input data and weights of the neural network and obtaining positional indices of non-zero weights; traversing a training set and obtaining multiple data sets related to the non-zero weights based on the positional indices of the non-zero weights using the incompletely specified function; performing logical minimization on the multiple data sets to generate a Boolean logic expression based on the incompletely specified function; determining the size and internal port connections of a flash memory logic array based on a cube of the Boolean logic expression and logical variables in the Boolean logic expression; constructing a flash memory logic array based on the size and internal port connections of the flash memory logic array, and constructing peripheral circuits for the flash memory logic array, wherein the peripheral circuits are primarily used to process the output data of the flash memory logic array. The present invention reduces power consumption and latency during neural network hardware implementation.
Owner:HUBEI UNIV

Hash-based sparse matrix vector multiplication optimization method and device

The application provides a hash-based sparse matrix vector multiplication optimization method, which is characterized by the following steps: dividing a sparse matrix to be multiplied into a plurality of sparse matrix blocks according to the hardware structure of a neural network hardware accelerator, performing linear hash mapping on the plurality of sparse matrix blocks to obtain a to-be-divided matrix; dividing the to-be-divided matrix into a plurality of sub-matrix blocks according to the size of the to-be-divided matrix and the hardware structure, and dividing parallel execution parts and competitive execution parts in the sub-matrix blocks; the neural network hardware accelerator performs competitive execution of a calculation task between the sub-matrix blocks and parallel execution of the calculation task in the sub-matrix blocks to obtain a plurality of sub-matrix calculation results, and restores the original order of writing through a hash table; and the plurality of sub-matrix calculation results are combined according to the original order to obtain a final result of matrix vector multiplication.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Self-adaptive non-volatile neural network circuit

The invention relates to the field of neural network hardware circuits, in particular to a self-adaptive non-volatile neural network circuit which comprises a plurality of memristor arrays and an adjusting circuit, the memristor arrays are used for storing weight information, processing input signals and outputting calculation results to the adjusting circuit for processing information output by the memristor arrays, and the adjusting circuit is used for adjusting the weight information of the memristor arrays. The adjusting circuit comprises a controller and a multiplexer, a redundant node optimization module is arranged in the controller, adaptive adjustment of the size and the resistance value of the memristor array is realized through the redundant node optimization module and the adjusting circuit, and secondary adjustment is performed on the array during on-chip training. Therefore, the neural network can adaptively and efficiently operate when processing different tasks. The neural network hardware circuit is low in power consumption, high in precision and high in adaptability.
Owner:SHANDONG UNIV

An in-situ compensation method for hardware accuracy issues in Skip Structure deep neural networks

This invention discloses an in-situ compensation method for the hardware accuracy problem of Skip Structure deep neural networks, belonging to the field of memristor technology. The in-situ compensation method includes S1, determining the correlation coefficient of the compensation equation; S2, testing the relationship between the correlation coefficient and the output error; S3, performing data fitting based on the test data obtained in S2, and establishing the compensation equation; S4, when the array outputs the results, using the compensation equation established in S3 to perform in-situ compensation on the output results. The compensation method of this invention can solve the accuracy degradation problem caused by the layer-by-layer accumulation of errors when implementing Skip Structure deep neural networks with memristor arrays, thus having a significant optimization effect on the hardware implementation of skip structure neural networks.
Owner:ANHUI UNIV

A 2-group signed tensor computing circuit structure based on 6-bit approximate full adder

The application discloses a 2-group signed tensor calculation circuit structure based on a 6-bit approximate full adder, relates to the field of neural network hardware acceleration, and comprises a 6-bit approximate full adder module, a signed 8*8 approximate multiplier circuit structure based on the 6-bit approximate full adder and the 2-group signed tensor calculation circuit structure based on the 6-bit approximate full adder. Since the neural network accelerator can sacrifice part of the accuracy of data in exchange for the optimization of delay, area and power consumption of a circuit structure, the 6-bit approximate full adder module ignores part of the carry, thereby reducing the circuit area and lowering the circuit power consumption, the signed 8*8 multiplier calculation process and the 2-group signed tensor calculation process are optimized by using the 6-bit approximate full adder module, approximate calculation is introduced at some positions, and in exchange for the improvement of area and power consumption of the circuit structure, part of the accuracy is lost.
Owner:SOUTHEAST UNIV

An electromagnetic echo signal processing method based on programmable electromagnetic neural network

The application discloses an electromagnetic echo signal processing method based on a programmable electromagnetic neural network, and the system comprises a vehicle, an obstacle, a transmitting end SPNN, a transmitting end antenna array, a receiving end antenna array, a receiving end module, a receiving end SPNN, an intensity detection module, a collection and decision module; the method is: based on the programmable electromagnetic neural network, a transmitting end electromagnetic neural network code is designed; the transmitting and receiving antenna arrays are installed, and echo data sets of the receiving antenna array under various obstacle scenes are collected; a receiving end electromagnetic neural network code and a back-end processing matrix are designed; output data sets of the output end programmable electromagnetic neural network are collected under actual working conditions; the back-end processing matrix is optimized again; the electromagnetic echo signal processing method has high-speed and high-precision rate identification ability for potential obstacle targets, can effectively improve the response rate, overcomes the hardware requirements of multi-port electromagnetic port receiving, and simultaneously meets the low-power consumption demand.
Owner:SOUTHEAST UNIV

Graphics architecture including a neural network pipeline

One embodiment provides a graphics processor comprising a block of graphics cores and circuitry including a programmable neural network unit, the programmable neural network unit including one or more neural network hardware blocks, wherein a neural network hardware block includes circuitry to perform neural network operations and activation operations for a layer of a neural network, the programmable neural network unit addressable by cores within the block of graphics cores, wherein the programmable neural network unit is to configure one or more neural network hardware blocks with a meta-shader neural network, the meta-shader neural network to generate a texture for one of multiple types of terrain.
Owner:INTEL CORP

High speed optical neural network hardware accelerator using adiabatic elimination-based ITO optical logic gates

A photonic gate system comprising a center waveguide that is provided with a continuous wave input; a first electrically controlled plasmonic waveguide configured on a first opposing side that is adjacent to the center waveguide; a second electrically controlled plasmonic waveguide configured on a second opposing side that is adjacent to the center waveguide; a first outer waveguide configured adjacent to the first electrically controlled plasmonic waveguide; and a second outer waveguide configured adjacent to the second electrically controlled plasmonic waveguide.
Owner:UNIV OF FLORIDA RESEARCH FOUNDATION INC

A heterogeneous control circuit based on spiking neural networks

The application belongs to the technical field of pulse neural network hardware, and particularly relates to a heterogeneous control circuit based on a pulse neural network. The application provides a heterogeneous control circuit based on a pulse neural network, which realizes expansion of coprocessor instructions through a high-universal peripheral interface and a sub-instruction receiving module, improves the overall design flexibility of the heterogeneous control circuit, and meanwhile, the application improves the parallel degree of overall instruction sending through a sub-instruction cache module and effectively solves the data conflict problem caused by the improvement of the parallel degree through a proposed instruction lock module.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Global modulo allocation in neural network compilation

In one example, a method performed by a compiler comprises: receiving a dataflow graph of a neural network, the neural network comprising a neural network operator; receiving information of computation resources and memory resources of a neural network hardware accelerator intended to execute the neural network operator; determining, based on the dataflow graph, iterations of an operation on elements of a tensor included in the neural network operator; determining, based on the information, a mapping between the elements of the tensor to addresses in the portion of the local memory, and a number of the iterations of the operation to be included in a batch, wherein the number of the iterations in the batch are to be executed in parallel by the neural network hardware accelerator; and generating a schedule of execution of the batches of the iterations of the operations.
Owner:AMAZON TECH INC