Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

13 results about "Neural network hardware" patented technology

Low-power convolutional neural network hardware acceleration method based on stochastic computing

This invention discloses a low-power convolutional neural network hardware acceleration method based on random computation. First, a convolutional neural network model is constructed and trained. Then, during the inference phase, the trained model is deployed to a hardware computing unit, and random computation is used to partially or completely replace the convolutional computations in the model. The random computation includes three stages: encoding, computation, and decoding. In the encoding stage, the input feature map is scaled, and the convolutional kernel is quantized. The gray values ​​of the convolutional kernel and each point in the input feature map are normalized into probability values, and an encoding sequence is generated for each probability value. In the computation stage, the two encoding sequences of the convolutional kernel and corresponding points in the input feature map are subjected to a traversal AND operation to generate an AND operation sequence. In the decoding stage, the AND operation sequence is decoded to obtain the computation result. Without significantly affecting model performance, this method reduces computational complexity and hardware resource consumption, making it suitable for scenarios with limited power consumption and resources.
Owner:HEBEI UNIV OF TECH

A 2-group signed tensor computing circuit structure based on 6-bit approximate full adder

ActiveCN115840556BBinary multiplierNeural network hardware
The application discloses a 2-group signed tensor calculation circuit structure based on a 6-bit approximate full adder, relates to the field of neural network hardware acceleration, and comprises a 6-bit approximate full adder module, a signed 8*8 approximate multiplier circuit structure based on the 6-bit approximate full adder and the 2-group signed tensor calculation circuit structure based on the 6-bit approximate full adder. Since the neural network accelerator can sacrifice part of the accuracy of data in exchange for the optimization of delay, area and power consumption of a circuit structure, the 6-bit approximate full adder module ignores part of the carry, thereby reducing the circuit area and lowering the circuit power consumption, the signed 8*8 multiplier calculation process and the 2-group signed tensor calculation process are optimized by using the 6-bit approximate full adder module, approximate calculation is introduced at some positions, and in exchange for the improvement of area and power consumption of the circuit structure, part of the accuracy is lost.
Owner:SOUTHEAST UNIV

An electromagnetic echo signal processing method based on programmable electromagnetic neural network

PendingCN122151029ABiological modelsRadio wave reradiation/reflectionData setNeural network hardware
The application discloses an electromagnetic echo signal processing method based on a programmable electromagnetic neural network, and the system comprises a vehicle, an obstacle, a transmitting end SPNN, a transmitting end antenna array, a receiving end antenna array, a receiving end module, a receiving end SPNN, an intensity detection module, a collection and decision module; the method is: based on the programmable electromagnetic neural network, a transmitting end electromagnetic neural network code is designed; the transmitting and receiving antenna arrays are installed, and echo data sets of the receiving antenna array under various obstacle scenes are collected; a receiving end electromagnetic neural network code and a back-end processing matrix are designed; output data sets of the output end programmable electromagnetic neural network are collected under actual working conditions; the back-end processing matrix is optimized again; the electromagnetic echo signal processing method has high-speed and high-precision rate identification ability for potential obstacle targets, can effectively improve the response rate, overcomes the hardware requirements of multi-port electromagnetic port receiving, and simultaneously meets the low-power consumption demand.
Owner:SOUTHEAST UNIV

A heterogeneous control circuit based on spiking neural networks

ActiveCN118194941BCoprocessorSpiking neural network
The application belongs to the technical field of pulse neural network hardware, and particularly relates to a heterogeneous control circuit based on a pulse neural network. The application provides a heterogeneous control circuit based on a pulse neural network, which realizes expansion of coprocessor instructions through a high-universal peripheral interface and a sub-instruction receiving module, improves the overall design flexibility of the heterogeneous control circuit, and meanwhile, the application improves the parallel degree of overall instruction sending through a sub-instruction cache module and effectively solves the data conflict problem caused by the improvement of the parallel degree through a proposed instruction lock module.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

A high-precision automated deployment method for analog computing-in-memory accelerator

PendingCN122452656ANeural network hardwareFloating point
The application relates to the technical field of neural network hardware acceleration, in particular to a high-precision automatic deployment method for an analog memory-computation integrated accelerator; the method comprises the following steps: layer-by-layer statistics of activation values and weight distribution of a floating-point model with a verification set is carried out, a quantization scale factor is automatically updated to avoid information truncation; quantization parameters are constrained in a discrete value set that can be realized by a chip, an interlayer mathematical consistency relationship is established, and a cross-layer linkage updating mechanism is adopted to process parameter coupling; based on an additive Gaussian noise model, weight and activation dynamic ranges are jointly optimized from the perspective of improving end-to-end signal-to-noise ratio, amplitude compensation and cross-layer correction are carried out; with deployment precision or signal-to-noise ratio as the target, a closed-loop iteration is constructed between discrete quantization training and joint compensation until the precision meets the standard or the parameters converge; a weight programming file, a quantization parameter configuration file and an inference control file are generated, which are directly loaded by the chip. The application realizes high-precision automatic deployment of the analog memory-computation integrated accelerator.
Owner:PAIFANG TECH (BEIJING) CO LTD

A cross-block continuous memory access scheduling method and a neural network hardware accelerator

PendingCN122434719AComputer hardwareNeural network hardware
The application discloses a cross-block continuous memory access scheduling method and a neural network hardware accelerator, and belongs to the technical field of memory access. The method comprises the following steps: dividing a to-be-processed image into a plurality of sequentially arranged image blocks; starting from the first image block, generating control information for at least one image block, and sequentially writing the control information into an information FIFO, and each time reading the next control information; repeatedly performing a read address request operation until all the read address requests of the image blocks are sent or a preset maximum Outstanding depth is reached; in parallel with the read address request operation, sequentially receiving data of each image block returned by a DDR according to the order in which the read address requests are sent, and whenever the data reception of an image block is completed, if the information FIFO is not empty, immediately starting to receive data of the next image block until the data of all the image blocks whose address requests have been sent are completely received. The application completely eliminates the idle period of inter-block memory access.
Owner:ZHEJIANG XINMAI SILICON CO LTD

In-memory computing method and device based on in-memory index and multi-layer perceptron

The application provides in-memory index-based multilayer perceptron high-energy-efficiency in-memory calculation, and relates to the technical field of neural network hardware acceleration. The method comprises the following steps: extracting weight parameters of each layer in a multilayer perceptron model, generating a lookup table, and storing the lookup table in on-chip memory; according to the network layer type, dynamically loading a weight sub-table from the lookup table into a multiple address concurrent cache to form a local lookup table for single-cycle multi-address parallel access; according to the network layer type and a block strategy, batch reading of input feature vectors is performed and the input feature vectors are divided into multiple data blocks; using an activation function, feature values in each data block are transformed into index addresses, the local lookup table is accessed in parallel, and multiple weight data are obtained as lookup table results in a single cycle; through a hierarchical reduction circuit, the lookup table results are accumulated and fused to generate an activation output value; the activation output value is format-converted and written back to a feature data storage as input features of the next layer network or a model inference result. In this way, the requirements of low power consumption and high throughput are effectively balanced.
Owner:XIDIAN UNIV

Fault-tolerant methods, systems, and media for neural network hardware inference in BNCT radiation environments

PendingCN122309231AComputer hardwareData stream
This invention discloses a fault-tolerant method, system, and medium for neural network hardware inference in BNCT radiation environments, relating to the fields of computer architecture and hardware fault-tolerant technology. The method includes: pre-dividing the weights of a 3D deep learning model into multiple weight blocks according to on-chip storage capacity and generating checksums; during the inference phase, dynamically loading the weights into the on-chip cache block by block and performing hardware-level integrity checks; when a radiation-induced check failure is detected, performing a local rollback operation for the erroneous weight block without clearing the already calculated intermediate feature maps or interrupting the current data flow inference pipeline. This invention enables fault-tolerant deployment of large models on hardware such as FPGAs, ensuring the continuity of neutron flux distribution prediction and radiation resistance safety.
Owner:SICHUAN UNIV

A Convolutional Neural Network Accelerator System Based on a RISC-V Processor

PendingCN122311316AComputer architectureNeural network hardware
A convolutional neural network accelerator system based on a RISC-V processor belongs to the field of neural network hardware acceleration technology. The RISC-V processor described in this invention is a Hummingbird E203 processor, and the convolutional neural network accelerator system is connected to the Hummingbird E203 processor via a NICE interface. The NICE interface is used to realize data interaction between the Hummingbird E203 processor and the convolutional neural network accelerator system, including the transmission of custom extended instructions, memory data, and result data. The convolutional neural network accelerator system includes a weight configuration module, a sliding window generation module, a convolution calculation module, an activation module, a pooling module, a controller, a decoder, a data input module, and a data output module. This invention is used to implement convolution calculation, activation processing, and pooling processing in convolutional neural networks, and to achieve collaborative control between the RISC-V processor and the accelerator.
Owner:JIANGNAN UNIV

Configurable convolution accelerator device supporting shared micro-exponential format data and method thereof

PendingCN122366551AVery large scale integrated circuitsZero padding
This invention discloses a configurable convolution acceleration device and method supporting shared micro-exponential format data, belonging to the field of neural network hardware acceleration for very large-scale integrated circuits. This acceleration device accelerates convolutional neural networks quantized using a shared micro-exponential data format. The device features a configurable design, supporting the following convolution modes: a kernel size of [value missing], a stride of 2, and zero padding of 1; and a kernel size of [value missing], a stride of 1, and zero padding of 0. The supported input feature map channels and the maximum number of rows and columns supported for these channels are: 32 / 320, 64 / 160, 128 / 80, and 256 / 40. Furthermore, the design method for this configurable convolution acceleration device, supporting new data formats, can also achieve convolution support for different MX quantization formats for input feature map data and weight data by modifying the design parameters.
Owner:NANJING UNIV

Alu operation fusion processing module and method suitable for neural network

The application discloses an ALU operation fusion processing module and method suitable for a neural network, and the module comprises: a control unit, which is used for receiving and decoding machine instructions, managing the execution flow of inner and outer loops and microinstruction loops, and generating control signals of each stage of a microinstruction pipeline; a microinstruction buffer, which is used for storing a microinstruction sequence generated in advance by a neural network compiler; a register file, which is used for storing source operands and results of ALU operations; each entry of the register file is composed of a valid bit, a tag bit and a data bit; an ALU calculation core, which adopts a SIMD architecture and comprises multiple parallel arithmetic logic function units; a Load / Store unit, which is used for processing data exchange between the register file and a local buffer; and a data selection interface, which is used for selecting a data source or a target buffer according to a storage tag field of a microinstruction. The application can effectively balance the efficiency and flexibility of the ALU unit in a neural network hardware accelerator.
Owner:ZHEJIANG UNIV

Hardware implementations of neural networks

Hardware implementations of neural networks and methods of processing data in such hardware implementations are disclosed. Input data for a plurality of layers of the network is processed in blocks to generate corresponding blocks of output data. The processing is performed through the plurality of layers in a depth direction, evaluating all of the layers for a given block before proceeding to the next block.
Owner:IMAGINATION TECH LTD