Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

61 results about "Neural network hardware" patented technology

Pulse neural network hardware accelerator and data processing method

The invention discloses a pulse neural network hardware accelerator and a data processing method, and the accelerator is characterized in that a low-power-consumption three-stage pipeline CPU module is used for receiving input data, scheduling an SNN network acceleration instruction, and sending the input data to an asynchronous edge SNN hardware accelerator module through a coprocessor interface; the asynchronous edge SNN hardware accelerator module comprises a pulse data encoding and decoding module, L neuromorphic kernels and an on-chip network, the pulse data encoding and decoding module encodes input data into a pulse form and sends the pulse form into the neuromorphic kernels, and the neuromorphic kernels are used for performing calculation based on the data in the pulse form; the network-on-chip is used for communication between the neuromorphic kernels, and the connection between neurons before and after synapses in the neuromorphic kernels is realized by adopting a synaptic cross array. According to the invention, the data processing acceleration performance can be greatly improved.
Owner:WUHAN UNIV +1

Graph Neural Network Hardware Accelerator

The description relates to graph neural network hardware accelerators. One example can include multiple FPGAs or ASICs that each include multiple parallel arranged processing elements and a shared memory. Individual processing elements are configured to prune a subgraph of a graph neural network model. The shared memory is configured to recombine the pruned subgraphs to generate a pruned graph neural network model.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Dynamic convolution calculation method and architecture for hardware acceleration

The invention discloses a dynamic convolution calculation method and architecture for hardware acceleration, relates to the technical field of neural network hardware acceleration, and solves the technical problem that the current calculation method and architecture urgently need a hardware architecture capable of flexibly adjusting a convolution kernel, stride and filling parameters. The method comprises the steps of dynamically adjusting convolution calculation parameters in a convolution calculation task according to input data; decomposing the convolution calculation task into a plurality of subtasks according to the availability of hardware resources and the calculation amount of each convolution operation in the convolution calculation task; distributing the plurality of sub-tasks to different computing units according to a priority scheduling strategy; monitoring the load condition of each calculation unit in real time, and carrying out load balancing on each calculation unit; and executing each sub-task after load balancing in parallel through a plurality of calculation units to obtain a calculation result of the convolution calculation task. According to the invention, convolution calculation parameters can be flexibly adjusted, and the calculation efficiency and the resource utilization rate of the NPU are improved.
Owner:林培东 +2

ALU operation fusion processing module and method suitable for neural network

The invention discloses an ALU operation fusion processing module and method suitable for a neural network, and the module comprises a control unit which is used for receiving and decoding a machine instruction, managing the execution processes of internal and external circulation and microinstruction circulation, and generating a control signal of each stage of a microinstruction assembly line; the microinstruction buffer area is used for storing a microinstruction sequence pre-generated by the neural network compiler; the register file is used for storing source operands and results of ALU operation; each entry of the register file is composed of a valid bit, a tag bit and a data bit; the ALU computing core adopts an SIMD (Single Instruction Multiple Data) architecture and comprises a plurality of paths of parallel arithmetic logic function units; the Load / Store unit is used for processing data exchange between a register file and a local buffer area; and the data selection interface is used for selecting a data source or a target buffer area according to the storage tag field of the microinstruction. According to the method, the high efficiency and the flexibility of the ALU in the neural network hardware accelerator can be effectively considered.
Owner:ZHEJIANG UNIV

Convolutional neural network hardware acceleration method and system

The invention relates to the technical field of hardware acceleration, and provides a convolutional neural network hardware acceleration method and system, and the method comprises the steps: obtaining a convolutional neural network instruction which comprises an img2col instruction, a systolic array multiplication instruction and a col2img instruction; analyzing an operation instruction code and a function code field of each instruction, determining a target register, a source register and calculation parameters, determining an instruction function and generating a control signal; data to be operated are taken out from the source register and transmitted to the corresponding operation module together with the control signal and the calculation parameter to be calculated, and then the data are written back to the target register; through hardware optimization strategies such as multi-level address superposition, systolic array data multiplexing, double-buffer window construction and a parallel comparison tree, extra power consumption caused by multiple times of data loading and storage and intermediate result calculation is reduced, and data handling overhead, calculation delay and control complexity are greatly reduced.
Owner:SHANDONG UNIV

Softmax instruction set extension method and system based on RISC-V

The invention belongs to the field of neural network hardware acceleration, and provides a Softmax instruction set extension method and system based on RISC-V. The Softmax instruction set extension method based on the RISC-V. The Softmax instruction set extension method based on the RISC-V. The Softmax instruction set extension method based on the RISC-V. The Softmax instruction set extension method based on the RISC-V. The Softmax instruction set extension method comprises an instruction fetching stage, in the decoding stage, corresponding instruction functions are analyzed for Opcode, Funct7 and Funct3 of the Softmax instruction, and corresponding control signals are generated. In the execution stage, data to be subjected to Softmax operation is taken out from the data memory and transmitted into the Softmax calculation unit according to a control signal; executing Softmax calculation according to the following formula, and writing a Softmax calculation result back to the target register; the calculation process is divided into four sub-modules of maximum solution, index calculation, summation and normalization, and a hardware acceleration strategy is designed, so that the operation delay and the resource overhead are greatly reduced, the degree of parallelism, the precision and the energy efficiency ratio of calculation are effectively improved, and the method is suitable for a high-performance neural network reasoning acceleration scene.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Streamline processing graph to enhance performance of neural network execution on hardware accelerator

Neural networks have interconnected operations that process data, and the layered structure of neural networks enable complex machine learning tasks. Neural network hardware accelerators with specialized circuitry can perform the operations in an efficient manner. Some neural network hardware accelerators have a programmable look up table that can be used to efficiently apply an approximation of an activation function to an input tensor. To further improve efficiency, the programmable look up table can be used to store parameters that can be used to approximate a composite function that combines two or more activation functions. In some cases, the programmable look up table can be used to store two or more sets of parameters that can be used to approximate different two or more activation functions.
Owner:INTEL CORP +7

A hardware accelerator and acceleration method for convolutional neural networks

This invention discloses a convolutional neural network hardware accelerator, comprising: a memory that stores input feature map data in an input channel-first order; a feature map input module that reads input feature map data in an input channel-first reading order; a feature map caching module that caches the input feature map data in an input channel-first storage order; a convolution kernel module that stores convolution kernel data in an input channel-first storage order; a computation module that sequentially and in parallel reads M input feature map data and M convolution kernel data in each clock cycle, performs point-to-point multiplication on the input feature map data and the M convolution kernel data, accumulates all the point-to-point multiplication results to obtain the convolution result, and outputs the convolution result in parallel to N output channels as output feature map data; and a feature map output module that writes the output feature map data into the memory in an output channel-first storage order. This invention improves the computational efficiency of convolution operations in the hardware accelerator.
Owner:HANGZHOU FEISHU TECH CO LTD

Neural network hardware device and system

Neural network systems and methods are provided. In one embodiment, a method of making a neural network device includes: forming a mesh layer on a substrate, the mesh layer including a matrix of randomly dispersed conductive nano-strands insulated from one another; forming an isolation trench extending into the mesh layer; forming a memristor device extending into the mesh layer, the memristor device including: an electrical conductor, and a layer of memristive material in electrical contact with individual nano-strands of a first set of conductive nano-strands in the mesh layer; forming an electrode extending into the mesh layer and spaced from the memristor device by the isolation trench, wherein the electrode is in electrical contact with individual conductive nano-strands of a second set of conductive nano-strands in the mesh layer; and forming a modulating device bridging the memristor device and the electrode.
Owner:THE GOVERNMENT OF THE UNITED STATES OF AMERICA AS REPRESENTED BY THE SECRETARY DEPARTMENT OF HEALTH & HUMAN SERVICES

Pipelined architecture based computing engine acceleration apparatus, method, device and medium

The application provides a pipeline structure-based computing engine acceleration device, method, equipment and medium in the technical field of neural network hardware acceleration, the device comprises: a circular row buffer, comprising: a first multiplexer, used for determining the storage position of n feature map pixels of an input feature map, wherein n is a positive integer greater than 0; m random access memories, used for caching n feature map pixels according to the storage position, wherein m is a positive integer greater than 0; a second multiplexer, used for outputting the feature map pixels written into the m random access memories; a stream driving filler, used for performing boundary filling on the feature map pixels output by the second multiplexer according to the position of the feature map pixels and the boundary parameter of the convolution layer in the neural network; a computing unit array, used for performing convolution calculation on the filled feature map pixels; and a weight memory, used for caching the weight data of the neural network.
Owner:INST OF SEMICONDUCTORS - CHINESE ACAD OF SCI

Hardware acceleration system and method supporting pulse-gated recurrent neural network

The invention discloses a hardware acceleration system and method supporting a pulse-gated recurrent neural network, and relates to the technical field of neural network hardware acceleration, and the system comprises a PC terminal and an accelerator. The accelerator comprises a UART (Universal Asynchronous Receiver / Transmitter) module, a data composer, a storage module, a controller module, a calculation module and a gating calculation core, the PC terminal is connected with the UART module; the gating calculation core is connected with the data composer; and the controller module and the storage module are respectively connected with the UART module, the data composer, the calculation module and the gating calculation core. Based on the above scheme, different coding modes can be selected at a PC terminal to adapt to different data types, and a pulse neuron activation function and a logic gate are utilized to perform calculation heterogeneous design in a calculation module and a gating calculation core, so that a high-performance configurable hardware acceleration calculation architecture for the pulse gating recurrent neural network is realized.
Owner:GUANGDONG UNIV OF TECH

Large language model softmax function hardware acceleration circuit and method

The invention discloses a large language model softmax function hardware acceleration circuit and method, and belongs to the field of neural network hardware acceleration of super-large scale integrated circuits. An input sequence is divided into a plurality of data blocks which are processed in parallel, and three-stage pipeline division is adopted, so that average single calculation delay is shortened to G clock cycles, the calculation parallelism is improved, the calculation speed of a softmax function is improved, and the reasoning delay is reduced; and a sparse mask strategy of sparse threshold comparison is introduced, the sparsity of data is fully utilized, the problems of high calculation complexity and high calculation delay of the softmax function and the data access bottleneck of the softmax function are solved, the calculation cost is remarkably reduced, and the calculation efficiency is improved. In addition, a softmax function hardware circuit adaptive to software optimization is constructed, calculation delay and memory access pressure are reduced through the hardware circuit, and data processing efficiency is improved.
Owner:NANJING UNIV

Low-power convolutional neural network hardware acceleration method based on stochastic computing

This invention discloses a low-power convolutional neural network hardware acceleration method based on random computation. First, a convolutional neural network model is constructed and trained. Then, during the inference phase, the trained model is deployed to a hardware computing unit, and random computation is used to partially or completely replace the convolutional computations in the model. The random computation includes three stages: encoding, computation, and decoding. In the encoding stage, the input feature map is scaled, and the convolutional kernel is quantized. The gray values ​​of the convolutional kernel and each point in the input feature map are normalized into probability values, and an encoding sequence is generated for each probability value. In the computation stage, the two encoding sequences of the convolutional kernel and corresponding points in the input feature map are subjected to a traversal AND operation to generate an AND operation sequence. In the decoding stage, the AND operation sequence is decoded to obtain the computation result. Without significantly affecting model performance, this method reduces computational complexity and hardware resource consumption, making it suitable for scenarios with limited power consumption and resources.
Owner:HEBEI UNIV OF TECH

A transconductance variable field effect transistor array and applications

The application discloses a transconductance variable field effect transistor array suitable for a dendritic network hardware and application, and belongs to the technical field of semiconductor integrated circuits.The application realizes three-element multiplication of a storage variable and two input variables based on a single transconductance variable field effect transistor, and realizes mapping of a dendritic network core algorithm based on a complementary device array.Compared with a traditional neural network hardware which realizes nonlinear transformation by using a neuron activation circuit, the application realizes nonlinear transformation by using intrinsic nonlinearity of a device, effectively reduces design complexity, optimizes area and power consumption of a system peripheral circuit, and has important significance for design of a high-performance artificial intelligence computing system.
Owner:PEKING UNIV

Hash-based sparse matrix vector multiplication optimization method and device

The application provides a hash-based sparse matrix vector multiplication optimization method, which is characterized by the following steps: dividing a sparse matrix to be multiplied into a plurality of sparse matrix blocks according to the hardware structure of a neural network hardware accelerator, performing linear hash mapping on the plurality of sparse matrix blocks to obtain a to-be-divided matrix; dividing the to-be-divided matrix into a plurality of sub-matrix blocks according to the size of the to-be-divided matrix and the hardware structure, and dividing parallel execution parts and competitive execution parts in the sub-matrix blocks; the neural network hardware accelerator performs competitive execution of a calculation task between the sub-matrix blocks and parallel execution of the calculation task in the sub-matrix blocks to obtain a plurality of sub-matrix calculation results, and restores the original order of writing through a hash table; and the plurality of sub-matrix calculation results are combined according to the original order to obtain a final result of matrix vector multiplication.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Self-adaptive non-volatile neural network circuit

The invention relates to the field of neural network hardware circuits, in particular to a self-adaptive non-volatile neural network circuit which comprises a plurality of memristor arrays and an adjusting circuit, the memristor arrays are used for storing weight information, processing input signals and outputting calculation results to the adjusting circuit for processing information output by the memristor arrays, and the adjusting circuit is used for adjusting the weight information of the memristor arrays. The adjusting circuit comprises a controller and a multiplexer, a redundant node optimization module is arranged in the controller, adaptive adjustment of the size and the resistance value of the memristor array is realized through the redundant node optimization module and the adjusting circuit, and secondary adjustment is performed on the array during on-chip training. Therefore, the neural network can adaptively and efficiently operate when processing different tasks. The neural network hardware circuit is low in power consumption, high in precision and high in adaptability.
Owner:SHANDONG UNIV

An in-situ compensation method for hardware accuracy issues in Skip Structure deep neural networks

This invention discloses an in-situ compensation method for the hardware accuracy problem of Skip Structure deep neural networks, belonging to the field of memristor technology. The in-situ compensation method includes S1, determining the correlation coefficient of the compensation equation; S2, testing the relationship between the correlation coefficient and the output error; S3, performing data fitting based on the test data obtained in S2, and establishing the compensation equation; S4, when the array outputs the results, using the compensation equation established in S3 to perform in-situ compensation on the output results. The compensation method of this invention can solve the accuracy degradation problem caused by the layer-by-layer accumulation of errors when implementing Skip Structure deep neural networks with memristor arrays, thus having a significant optimization effect on the hardware implementation of skip structure neural networks.
Owner:ANHUI UNIV

A 2-group signed tensor computing circuit structure based on 6-bit approximate full adder

The application discloses a 2-group signed tensor calculation circuit structure based on a 6-bit approximate full adder, relates to the field of neural network hardware acceleration, and comprises a 6-bit approximate full adder module, a signed 8*8 approximate multiplier circuit structure based on the 6-bit approximate full adder and the 2-group signed tensor calculation circuit structure based on the 6-bit approximate full adder. Since the neural network accelerator can sacrifice part of the accuracy of data in exchange for the optimization of delay, area and power consumption of a circuit structure, the 6-bit approximate full adder module ignores part of the carry, thereby reducing the circuit area and lowering the circuit power consumption, the signed 8*8 multiplier calculation process and the 2-group signed tensor calculation process are optimized by using the 6-bit approximate full adder module, approximate calculation is introduced at some positions, and in exchange for the improvement of area and power consumption of the circuit structure, part of the accuracy is lost.
Owner:SOUTHEAST UNIV

An electromagnetic echo signal processing method based on programmable electromagnetic neural network

The application discloses an electromagnetic echo signal processing method based on a programmable electromagnetic neural network, and the system comprises a vehicle, an obstacle, a transmitting end SPNN, a transmitting end antenna array, a receiving end antenna array, a receiving end module, a receiving end SPNN, an intensity detection module, a collection and decision module; the method is: based on the programmable electromagnetic neural network, a transmitting end electromagnetic neural network code is designed; the transmitting and receiving antenna arrays are installed, and echo data sets of the receiving antenna array under various obstacle scenes are collected; a receiving end electromagnetic neural network code and a back-end processing matrix are designed; output data sets of the output end programmable electromagnetic neural network are collected under actual working conditions; the back-end processing matrix is optimized again; the electromagnetic echo signal processing method has high-speed and high-precision rate identification ability for potential obstacle targets, can effectively improve the response rate, overcomes the hardware requirements of multi-port electromagnetic port receiving, and simultaneously meets the low-power consumption demand.
Owner:SOUTHEAST UNIV

High speed optical neural network hardware accelerator using adiabatic elimination-based ITO optical logic gates

A photonic gate system comprising a center waveguide that is provided with a continuous wave input; a first electrically controlled plasmonic waveguide configured on a first opposing side that is adjacent to the center waveguide; a second electrically controlled plasmonic waveguide configured on a second opposing side that is adjacent to the center waveguide; a first outer waveguide configured adjacent to the first electrically controlled plasmonic waveguide; and a second outer waveguide configured adjacent to the second electrically controlled plasmonic waveguide.
Owner:UNIV OF FLORIDA RESEARCH FOUNDATION INC

A heterogeneous control circuit based on spiking neural networks

The application belongs to the technical field of pulse neural network hardware, and particularly relates to a heterogeneous control circuit based on a pulse neural network. The application provides a heterogeneous control circuit based on a pulse neural network, which realizes expansion of coprocessor instructions through a high-universal peripheral interface and a sub-instruction receiving module, improves the overall design flexibility of the heterogeneous control circuit, and meanwhile, the application improves the parallel degree of overall instruction sending through a sub-instruction cache module and effectively solves the data conflict problem caused by the improvement of the parallel degree through a proposed instruction lock module.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Implementation method of dynamic precision approximate multiplication based on partial product decoupling and multiplier

The invention provides a partial product decoupling-based dynamic precision approximate multiplication implementation method and a multiplier, and the multiplier comprises a significance evaluation unit which splits a first operand and a second operand into high digits and low digits, and correspondingly generates a first amplitude and a second amplitude; different precision instructions are generated based on the amplitudes and the threshold values; the dynamic precision unit performs zero setting processing on the second low digit based on the precision instruction to form an approximate number, and calculates a partial product of the first high digit and the approximate number; the accurate calculation unit accurately calculates products of the other three parts; and the final addition unit sums the products of the four parts according to the weight. The calculation unit can be used for constructing or optimizing a multiplication calculation array in a neural network hardware accelerator, is particularly suitable for scenes with dense fixed-point number operation such as convolution, and performs efficient approximate operation on the input fixed-point number through a data driving mechanism by utilizing the common sparsity and amplitude difference characteristics of neural network data, so that the calculation efficiency is improved. And self-adaptive management of power consumption is realized.
Owner:EHIWAY MICROELECTRONIC SCI & TECH (SUZHOU) CO LTD

A high-precision automated deployment method for analog computing-in-memory accelerator

The application relates to the technical field of neural network hardware acceleration, in particular to a high-precision automatic deployment method for an analog memory-computation integrated accelerator; the method comprises the following steps: layer-by-layer statistics of activation values and weight distribution of a floating-point model with a verification set is carried out, a quantization scale factor is automatically updated to avoid information truncation; quantization parameters are constrained in a discrete value set that can be realized by a chip, an interlayer mathematical consistency relationship is established, and a cross-layer linkage updating mechanism is adopted to process parameter coupling; based on an additive Gaussian noise model, weight and activation dynamic ranges are jointly optimized from the perspective of improving end-to-end signal-to-noise ratio, amplitude compensation and cross-layer correction are carried out; with deployment precision or signal-to-noise ratio as the target, a closed-loop iteration is constructed between discrete quantization training and joint compensation until the precision meets the standard or the parameters converge; a weight programming file, a quantization parameter configuration file and an inference control file are generated, which are directly loaded by the chip. The application realizes high-precision automatic deployment of the analog memory-computation integrated accelerator.
Owner:PAIFANG TECH (BEIJING) CO LTD

A cross-block continuous memory access scheduling method and a neural network hardware accelerator

The application discloses a cross-block continuous memory access scheduling method and a neural network hardware accelerator, and belongs to the technical field of memory access. The method comprises the following steps: dividing a to-be-processed image into a plurality of sequentially arranged image blocks; starting from the first image block, generating control information for at least one image block, and sequentially writing the control information into an information FIFO, and each time reading the next control information; repeatedly performing a read address request operation until all the read address requests of the image blocks are sent or a preset maximum Outstanding depth is reached; in parallel with the read address request operation, sequentially receiving data of each image block returned by a DDR according to the order in which the read address requests are sent, and whenever the data reception of an image block is completed, if the information FIFO is not empty, immediately starting to receive data of the next image block until the data of all the image blocks whose address requests have been sent are completely received. The application completely eliminates the idle period of inter-block memory access.
Owner:ZHEJIANG XINMAI SILICON CO LTD

In-memory computing method and device based on in-memory index and multi-layer perceptron

The application provides in-memory index-based multilayer perceptron high-energy-efficiency in-memory calculation, and relates to the technical field of neural network hardware acceleration. The method comprises the following steps: extracting weight parameters of each layer in a multilayer perceptron model, generating a lookup table, and storing the lookup table in on-chip memory; according to the network layer type, dynamically loading a weight sub-table from the lookup table into a multiple address concurrent cache to form a local lookup table for single-cycle multi-address parallel access; according to the network layer type and a block strategy, batch reading of input feature vectors is performed and the input feature vectors are divided into multiple data blocks; using an activation function, feature values in each data block are transformed into index addresses, the local lookup table is accessed in parallel, and multiple weight data are obtained as lookup table results in a single cycle; through a hierarchical reduction circuit, the lookup table results are accumulated and fused to generate an activation output value; the activation output value is format-converted and written back to a feature data storage as input features of the next layer network or a model inference result. In this way, the requirements of low power consumption and high throughput are effectively balanced.
Owner:XIDIAN UNIV

An improved neural network hardware acceleration method and device based on FPGA

The application relates to the technical field of image processing, in particular to an improved neural network hardware acceleration method and device based on FPGA, which comprises the following steps: training an SSD_MobilenetV1 network through transfer learning, data enhancement, multi-scale training and cosine annealing; performing structural pruning on the trained SSD_MobilenetV1 network, taking a convolution kernel or each network layer as a basic unit for pruning; adopting a QAT algorithm, introducing a pseudo-quantization operation for training, and using the pseudo-quantization operation to simulate the error of a quantization process; and converting the quantized SSD_MobilenetV1 network into a calculation graph. The application simultaneously uses an FPGA and an ARM processor to perform reasoning on a model, executes time-consuming convolution operators in the convolution network model in the FPGA, and executes other operators in the ARM processor, so that fast reasoning of the network model can be realized, power consumption is low, and the network model is beneficial to deployment to a terminal.
Owner:JIANGSU SIYUAN INTEGRATED CIRCUIT & INTELLIGENT TECH RES INST CO LTD

Data storage and access method and system for pulse convolutional neural network accelerator

The invention belongs to the technical field of neural network hardware accelerator and integrated circuit design. The invention provides a pulse convolutional neural network accelerator-oriented data storage and access method and system. According to the embodiment of the invention, a data storage and access mode matched with the systolic array calculation data flow is designed according to the binary property and the multi-time-step reasoning characteristic of the pulse data in the pulse convolutional neural network. A storage sequence which is mainly based on the input pulse data row direction, is grouped according to input channels and is continuously organized in the time step dimension is adopted in the off-chip memory, so that interaction with the off-chip memory can be carried out in a continuous access mode in the data loading and calculation result write-back process. The method comprises the following steps of: aggregating pulse data of a plurality of input channels in a storage word on an on-chip memory level; and meanwhile, in combination with an input cache design of a multi-BANK structure, multi-channel input data of adjacent rows can be read in parallel in a convolution window expansion process.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Data online compression method and device of neural network hardware accelerator

The application discloses a data online compression method and device of a neural network hardware accelerator. The method comprises converting a first activation value output by a neural network to obtain a first activation mask; dividing the first activation mask into at least two groups of activation sub-masks, and sequentially performing accumulation processing on each group of activation sub-masks in a preset order to obtain an activation position mask; calculating an activation selection mask based on the first activation mask, the activation position mask and a weight value output by the neural network; performing screening processing on the first activation value according to the activation selection mask to obtain a target activation value, and generating a second activation mask based on the target activation value. Through the online mask setting of the activation value and the offline compression of the weight value, the adaptability to different neural network compressions is high, the data moving efficiency can be improved, the power consumption is reduced, and the throughput is ensured.
Owner:JIANGNAN INST OF COMPUTING TECH

Compact configuration descriptor format for neural network hardware accelerator

PCT designated stageWO2026000671A1Multiprogramming arrangementsPhysical realisationMixture of expertsAlgorithm
Data processing units of a neural network hardware accelerator may be configured to perform neural network operations using configuration descriptors generated by a compiler. Growing sizes of neural network models can mean that the compiled configuration descriptors can become larger in size as well. Compile time can be longer, and storage overhead also increases. To address this issue, a compact configuration descriptor format having a constant configuration descriptor and a variable configuration descriptor can be used. The compact configuration descriptor format facilitates reuse across blocks, within blocks, and within a mixture of experts transformer-based machine learning model. The compact configuration descriptor format can dramatically reduce the size of the compiled configuration descriptors.
Owner:INTEL CORP +5

Circuit, method and neural network hardware system for implementing non-monotonic activation function

The application discloses a kind of circuit, method and neural network hardware system for realizing non-monotonic activation function, belong to the technical field of neural network hardware circuit design, the circuit includes: first inverter, second inverter and logic processing unit;The input end of first inverter and second inverter is connected, as analog accumulation signal input end;First inverter and second inverter are respectively set first logic flip threshold voltage and second logic flip threshold voltage of unequal value;The two input ends of logic processing unit are respectively connected to the output end of first inverter and the output end of second inverter.The application realizes non-monotonic activation function under extremely simple circuit architecture by the combination of double threshold decision and logic processing, without analog-digital conversion module, with the advantages of less transistor number, low power consumption, easy to integrate.
Owner:SOUTHEAST UNIV +1