Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

30 results about "Web accelerator" patented technology

A web accelerator is a proxy server that reduces web site access time. They can be a self-contained hardware appliance or installable software. Web accelerators may be installed on the client computer or mobile device, on ISP servers, on the server computer/network, or a combination. Accelerating delivery through compression requires some type of host-based server to collect, compress and then deliver content to a client computer.

Hybrid network-on-chip (NOC) for thread synchronization in many-core neural network accelerators

This application describes a network-on-chip system that could be used in a hardware accelerator for accelerating neural network computations. An example NoC system may include a plurality of cores, and a plurality of data links connecting adjacent cores of the plurality of cores for transmitting data. The example NoC may further include a global synchronization switch connected to each of the plurality of cores. The global synchronization switch is configured to dynamically connect any pair of cores in the plurality of cores for transmitting control signals between the pair of cores.
Owner:MOFFETT TECH CO LTD

MAC unit supporting input activation and weight double-end sparsity

The invention discloses an MAC unit supporting input activation and weight double-end sparsity, and relates to the technical field of neural network accelerator hardware design. Aiming at the technical defects existing in a sparse matrix acceleration scheme, the scheme is adopted, an activation value and non-zero elements of a weight matrix are stored through a special compressed data register, corresponding zero value distribution is recorded through a matched sparse bitmap register, and a joint sparse bitmap marking a double-non-zero effective calculation position is generated through dual-port joint operation; and the control and address generation subunit scans the joint sparse bitmap and outputs an address and an enable signal, the MAC calculation subunit is driven to read non-zero data to complete multiply-accumulate operation, and finally, an operation result is compressed and coded through an output end sparse encoder, so that efficient transmission of interlayer sparse data streams is realized. According to the method, input activation and weight double-end sparsity can be utilized at the same time, end-to-end sparse data flow is supported, zero value calculation can be dynamically skipped, and the energy efficiency ratio and the calculation throughput of neural network reasoning are improved.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

System and architecture of pure functional neural network accelerator

An accelerator circuit includes a control interface to receive a stream of instructions, a first memory to store an input data, and an engine circuit. The engine circuit includes a dispatch circuit to decode an instruction of the stream of instructions into a plurality of commands and a plurality of queue circuits. Each of the plurality of queue circuits supports a queue data structure to store a respective one of the plurality of commands decoded from the instruction, and a plurality of command execution circuits. Each of the plurality of command execution circuits is to receive and execute a command extracted from a corresponding one of the plurality of queues.
Owner:HUAXIA GENERAL PROCESSOR TECH INC

Efficient data movement for ai accelerators

Efficient data movement in neural network accelerators operating within virtualized memory systems is challenged by high address translation latency and unique data access patterns. To address this challenge, address translation prefetch (ATP) mechanisms can be implemented to proactively translate virtual memory addresses before data movement. ATP can be performed in advance of any data movement or concurrently with data movement while being throttled by page transition in the data movement request stream. The ATP mechanism can enforce quotas on outstanding ATP requests, independently for read and write streams, to preserve resources for other processes running on the neural network accelerator. In dealing with competing ATP requests, the mechanism can employ weighted arbitration to balance between different types of ATP requests, utilizing a programmable ratio. The ATP mechanisms enable scalable, high-throughput neural network inference in virtualized environments, addressing data movement bottlenecks in neural network accelerator deployments.
Owner:INTEL CORP

Hardware architecture, computing method and device for 3D point cloud neural network algorithm

The application relates to the technical field of point cloud neural network, in particular to a hardware architecture, a calculation method and equipment for a 3D point cloud neural network algorithm, wherein the hardware architecture comprises: off-chip storage; a mapping module, the mapping module is provided with distance filtering technology and / or output priority mapping calculation technology to reduce the off-chip memory access amount of the off-chip storage when calculating the mapping operation of the 3D point cloud neural network algorithm, and a mapping relationship between input and output is generated according to the mapping operation; and a calculation module, the calculation module is provided with an elastic array architecture, the elastic array architecture is adjusted according to different scales of calculation tasks in the 3D point cloud neural network algorithm, and the weight and the corresponding input feature are taken out according to the mapping relationship to perform matrix operation to obtain the output feature. Thus, the problems in the prior art that the off-chip memory access amount of the point cloud neural network accelerator is large, the utilization rate of the calculation unit is low, the processing speed of the accelerator is slow, the expansibility and flexibility are poor, and the actual needs cannot be met are solved.
Owner:TSINGHUA UNIVERSITY

Low-computational complexity neural network accelerator based on superposed pilot frequency

The invention relates to the technical field of wireless communication, in particular to a hardware accelerator of a channel estimation model based on a neural network in the field of wireless communication, the accelerator adopts a software and hardware collaborative design, and the core architecture comprises a control processing unit, a computing unit, an on-chip memory, an off-chip memory and an AXI interface, the control processing unit is integrated with a least square method module, a control module and a hardware interface module, the control module adopts a ping-pong buffer mechanism to schedule tasks, the computing unit comprises an OFDM decoder, a fixed-point quantization module, a superposition pilot frequency module, a channel judgment module, a neural network module and an output module, and by utilizing a superposition pilot frequency technology, the channel judgment module and the neural network module are integrated into the control processing unit. Through low-bit wide fixed-point quantization, sparse weight decoding and a special pipeline computing architecture, a computing link is constructed in a neural network module, and the scheme is used for solving the problem of high computing complexity when a deep learning channel estimation method is deployed in a resource-limited edge scene in the prior art.
Owner:CHANGCHUN UNIV OF SCI & TECH

Neural network accelerators, neural network acceleration methods, and devices

ActiveCN117273094BData segmentTerm memory
This invention relates to the field of artificial intelligence chips, providing a neural network accelerator and a method and apparatus for accelerating neural networks. The neural network accelerator includes: a direct memory access controller that continuously reads target data stored in external memory according to data arrangement order and sends it to a data warping and control module; the data warping and control module that generates at least one target data segment based on the received data according to its bit width, the bit width of the computing unit, and the convolution parameters; an accelerated computing module that performs accelerated computing on the target data segment from the data warping and control module based on at least some computing units, and outputs the computing results of the computing units to a data rearrangement module; and a data rearrangement module that merges the computing results of each computing unit according to the data arrangement order required for the accelerated computing of the next layer of the neural network as the input for the accelerated computing of the next layer. This improves acceleration performance.
Owner:INST OF SEMICONDUCTORS - CHINESE ACAD OF SCI

Reconfigurable hardware buffers in neural network accelerator framework

Embodiments of the present disclosure relate to reconfigurable hardware buffers in a neural network accelerator framework. A convolution accelerator framework (CAF) has a plurality of processing circuits including one or more convolution accelerators, a reconfigurable hardware buffer configurable to store data of a variable number of input data lanes, and a flow switch coupled to the plurality of processing circuits. The reconfigurable hardware buffer has a memory and control circuitry. The number of the variable number of input data lanes is associated with an execution epoch. During processing of the execution epoch, the flow switch streams data of the variable number of input data lanes between a processing circuit of the plurality of processing circuits and the reconfigurable hardware buffer. The control circuitry of the reconfigurable hardware buffer configures the memory to store the data of the variable number of input data lanes, the configuration including allocating a portion of the memory to each of the variable number of input data lanes.
Owner:STMICROELECTRONICS SRL

Filtering method, device and equipment of convolution-based FIR (Finite Impulse Response) digital filter and medium

The invention relates to the technical field of FIR digital filters, and discloses a filtering method, device and equipment of a convolution-based FIR digital filter and a medium, and the method comprises the steps: obtaining a one-dimensional convolution kernel and activation data corresponding to an N-order FIR digital filter; through LineBuf, filling operation is carried out on the activated data to overcome the boundary effect, and filled sub-data is obtained; and performing convolution operation based on the convolution kernel and the subdata through the MAC array to obtain a filtering result. The method is based on the convolution method and designs the LineBuf module loaded in stages, and benefits from the advantage, and a large amount of resource overhead can be reduced under the degree of parallelism. Besides, benefited from the advantages of the convolution method, the calculation parallelism degree is improved under the condition that the register overhead is not greatly increased, delay caused by software processing data filling is avoided through the LineBuf module, and high-speed calculation can be carried out under the condition of low resources. And meanwhile, the neural network accelerator and a common neural network accelerator are fused and share one MAC array, so that the hardware overhead is further reduced.
Owner:SHENZHEN UNIVERSITY OF ADVANCED TECHNOLOGY

Page fault support for virtual machine network accelerators

ActiveUS12712829B2Data packEngineering
Systems and methods for supporting page faults for virtual machine network accelerators. In one implementation, a processing device may receive, at a network accelerator device of a computer system, from a network, a first incoming packet and a second incoming packet. Responsive to receiving a first notification that an attempt to store the first incoming packet at a first buffer of a plurality of buffers associated with the network accelerator device caused a page fault, the processing device may store the first incoming packet at a second buffer and append a first identifier of the first buffer to a faulty buffer data structure. Responsive to receiving a second notification indicating a resolution of the page fault, the processing device may remove the first identifier from the faulty buffer data structure. The processing device may store the second incoming packet at the first buffer. The processing device may forward, to a driver of the network accelerator device, a second identifier of the second buffer and the first identifier of the first buffer.
Owner:RED HAT INC

Resource management method and heterogeneous chip

The invention relates to a resource management method, a heterogeneous chip and computer equipment. The method is applied to a bridging unit connected between a processor unit and a neural network accelerator, and comprises the following steps: acquiring physical state data and calculation state data of the neural network accelerator; according to the task queue data, priority arbitration is carried out on tasks of a to-be-executed task queue of the neural network accelerator, an arbitration result is obtained, and the arbitration result comprises existence of high-priority hard real-time tasks or absence of the high-priority hard real-time tasks; calculating state data and an arbitration result according to the physical state data, and generating a voltage regulation signal and a frequency regulation signal; and adjusting the voltage, the frequency and the clock switching state of the neural network accelerator according to the voltage adjusting signal and the frequency adjusting signal. By adopting the method, the energy efficiency, the real-time performance and the management uniformity of the heterogeneous computing system can be improved.
Owner:CCORE TECH CO LTD

Hardware queue manager serving multiple processor cores

PendingCN122663559AEngineeringData memory
A hardware-based queue manager (100) is implemented in a network accelerator (500). The queue manager can receive input from a processor core into a register set designated for the processor core. Other register sets can be designated for other processor cores. The queue manager reads the input in the register set, which triggers the queue manager to start a transaction, such as reading or writing to a queue data structure in a data memory (114). The queue data structure can handle multiple sizes of queue data structures.
Owner:TEXAS INSTRUMENTS INC

Accelerator device with automatic precision switching based on the sparsity rate

Device for automatically adjusting the arithmetic precision in a neural network accelerator, comprising: a) at least one computing unit (2) with a plurality of processing elements (1) supporting at least two precision modes (INT4, INT8, INT16); b) a sparsity measurement circuit in each computing unit which continuously measures the rate of skipped zero operations (4) in relation to the total number of operations (3) during the execution of multiply-accumulate operations within a configurable measurement window; c) a precision controller in each computing unit which automatically switches the arithmetic precision between the precision modes based on the measured sparsity rate (5);andd) a hysteresis logic in the precision controller with separate thresholds for switching up to lower precision (8, 10) and switching down to higher precision (9, 11), where TH_UP > TH_DOWN, to prevent oscillation at borderline sparsity rate.;
Owner:NEUROVEXON UG (HAFTUNGSBESCHRÄNKT)

Processing method and processing system utilizing a multi-core neural network accelerator

The invention is a processing method for carrying out a first processing task flow and at least one further processing task flow on a multi-core neural network accelerator having processing units (11, 12), a command scheduler, and a memory, wherein each processing task flow comprises commands (20, 21, 22, 23) and dependencies between at least some of the commands (20, 21, 22, 23). Each command (20, 21, 22, 23) is executed by one particular processing unit (11, 12). The scheduling of the commands (20, 21, 22, 23) is determined by a merged command schedule, maintaining the dependencies within the respective processing task flows and reducing idle times of the processing units (11, 12). The merged command schedule is determined by a command schedule merger unit based on the input command schedules and their timelines. The invention also relates to a processing system for carrying out the method.
Owner:AIMOTIVE KFT

Neural network accelerators performing operations with mixed format weights

The invention relates to a neural network accelerator performing operations with mixed format weights. The data processing unit may include a memory, a processing element (PE), and a control unit. The memory may store a weight block in a weight tensor of a neural network operation. Each weight block has an input channel (IC) dimension and an output channel (OC) dimension, and includes sub-blocks. The sub-block includes one or more weights having a first data precision and one or more other weights having a second data precision. The second data precision is lower than the first data precision. The control unit may distribute different sub-blocks to different PEs. The PE may receive the sub-block and perform a first MAC operation on a weight having a first data precision and a second MAC operation on a weight having a second data precision. The first MAC operation may consume more computation cycles or more multipliers than the second MAC operation.
Owner:ALTERA CORP

Cyclic Redundant Spare Testing (CREST)

PendingUS20260188414A1MultiplexingMultiplexer
A fault-tolerant computing architecture performs cyclic runtime testing and repair in column-oriented processing arrays. Operational columns are periodically paired with test columns that receive identical inputs, and their outputs are compared to detect faults while computation proceeds at full throughput. Upon detecting a validated mismatch, control circuitry substitutes a spare column for the faulty column. Inter-column multiplexers enable defect-location at a segment granularity by redirecting partial outputs to subsequent columns during diagnostic steps, after which normal routing is restored with the spare in place. The approach supports threshold-based validation to filter transient errors, pre-validation of spares, and scheduling at layer boundaries to avoid latency impact. The system enhances reliability of large neural-network accelerators without downtime and enables graceful degradation through dynamic substitution within the array.
Owner:SILVEBROOK KIA

Large language model operator conversion system and method and AI accelerator

The invention provides a large language model operator conversion system and method and an AI accelerator, the system comprises an operator identification module, a conversion module and a model reconstruction module, the operator identification module analyzes a computational graph structure of an input model to identify convertible operators meeting preset conversion conditions in the computational graph structure; the conversion module converts the convertible operator into an equivalent operator sequence based on a preset conversion strategy; and the model reconstruction module replaces nodes corresponding to convertible operators in the computational graph of the input model with an equivalent operator sequence, updates the connection relation and the weight initializer, and generates an updated model. According to the system, pipeline processing of model optimization is realized through an identification rule and a conversion strategy, and adaptation for different hardware platforms can be realized by updating a preset conversion condition and a conversion strategy, so that the problems of low deployment efficiency and low expandability on a CNN (Convolutional Neural Network) accelerator are solved.
Owner:NANJING UNIV

Network architecture, corresponding vehicle and method

The bandwidth of SOC interfaces is exploited while minimizing the number of physical ports via a networking accelerator for use on board a vehicle, for instance, that comprises: media access control (MAC) controller circuitry configured to provide a MAC port layer to control exchange of information, wherein the exchange of information comprises data flow transmission to virtual machine ports (VMPs) over a data link; virtual machine transmission (VM Tx) bridge circuitry configured to handle transmission data flow to the VMPs; transmission router / switch circuitry configured to route / switch data flow from the MAC controller circuitry to the VM Tx bridge circuitry; and queue handler circuitry configured to provide queue management for data flow between the MAC controller circuitry and the VM Tx bridge circuitry. The VM Tx bridge circuitry comprises virtual destination address circuitry configured to implement router / switch virtualization in the transmission router / switch circuitry with a virtual machine transmission descriptor based on a combination of a virtual machine port (VMP) tag indicative of a physical resource in the queue handler circuitry selectable for data flow transmission, and a virtual machine extended identifier (VMEID).
Owner:STMICROELECTRONICS INT NV

FPGA-based u-net network accelerator

The application discloses a U-Net network accelerator based on FPGA, and belongs to the fields of embedded systems and signal processing; a DDR interface is connected with an input cache module and a weight FIFO module through an input module; the input cache module and the weight FIFO module are connected with a convolution pooling module, and are used for storing feature map data required by current calculation; the convolution pooling module is connected with an output module through an output cache module; the output module is connected with the DDR interface, and is used for storing results obtained by convolution or pooling calculation; and the weight FIFO module is used for storing weight parameters required by current calculation; the application improves the utilization rate of internal hardware resources of FPGA, increases the parallel degree of the convolution operation hardware accelerator, and improves the overall operation performance of the hardware system.
Owner:HARBIN UNIV OF SCI & TECH

Hybrid network-on-chip (NOC) for thread synchronization in many-core neural network accelerators

This application describes a network-on-chip system that could be used in a hardware accelerator for accelerating neural network computations. An example NoC system may include a plurality of cores, and a plurality of data links connecting adjacent cores of the plurality of cores for transmitting data. The example NoC may further include a global synchronization switch connected to each of the plurality of cores. The global synchronization switch is configured to dynamically connect any pair of cores in the plurality of cores for transmitting control signals between the pair of cores.
Owner:MOFFETT TECH CO LTD

Instruction generation method and apparatus, pooling operation method, and electronic device

Disclosed are an instruction generation method and device, a pooling operation method and an electronic device, and relate to the technical field of neural networks. The method comprises determining hardware parameters of a neural network accelerator performing a pooling operation; determining pooling parameters of a pooling layer performing the pooling operation; the pooling parameters comprise a first pooling step, a size of an input feature map of the pooling layer, a size of a pooling kernel, and a type of the pooling operation; and generating instructions executable by the neural network accelerator according to the pooling parameters and the hardware parameters. The technical scheme of the present disclosure considers both the hardware parameters of the neural network accelerator and the pooling parameters of the pooling layer when generating the instructions executable by the neural network accelerator, so that the neural network accelerator can support pooling of multiple steps.
Owner:BEIJING HORIZON INFORMATION TECH CO LTD

Low power generative adversarial network accelerator and mixed-signal time-domain MAC array

Systems and methods for a low-cost mixed-signal time-domain accelerator for generative adversarial network (GAN) are provided. In one aspect, a system includes a memory and a training management unit (TMU) in communication with the memory. The TMU is configured to manage a training sequence. The system includes a time-domain multiplication-accumulation (TDMAC) unit in communication with the TMU, wherein the TDMAC unit is configured to perform time-domain multiplier operations and time-domain accumulator operations.
Owner:NORTHWESTERN UNIV

A neural network accelerator model conversion method and apparatus

The application discloses a neural network accelerator model conversion method and device, the method comprises the following steps: obtaining a neural network model to be converted, analyzing a model network structure file to obtain all network layers of the model, reconstructing the network layers, mapping the network layers into operator nodes supported by a neural network accelerator, and finally serializing the converted operator nodes and model weights according to a network topology structure to generate a target file; the device comprises a neural network model construction module, a reconstruction module, a mapping module and a serialization module; the application solves the multi-adaptation difficulty problem of deploying multiple format models on a neural network accelerator device, can efficiently convert the models, and generates a model format suitable for the neural network accelerator.
Owner:ZHEJIANG LAB +1

Hardware device for sparse attention calculation in neural networks

UndeterminedDE202026002052U1Physical realisationAlgorithmOperand
Device for processing attention scores in a hardware neural network accelerator, comprising: a) a compute array with a plurality of compute units (1), each compute unit having zero-detection logic which, on each computation operation, generates a sparsity flag (2) indicating whether at least one of the operands has the value zero; b) an attention unit with a score input which, in addition to a score value, has a dedicated port for the sparsity flag (3), the attention unit having a score buffer (4) in which received score values ​​are stored together with the associated sparsity flags;c) a hardware filter stage (8) arranged as a separate state in a state machine of the attention unit between score reception and softmax computation, wherein the filter stage iteratively traverses the stored scores and transfers non-null scores to a compacted buffer (5) with a reduced length (6); and d) a softmax pipeline operating exclusively on the compacted buffer (5) with the reduced length (6), thereby reducing the number of computational operations in the scaling, maximum search, exponential computation, and normalization stages proportionally to the sparsity rate.
Owner:NEUROVEXON UG (HAFTUNGSBESCHRÄNKT)

Hybrid network-on-chip (NoC) for thread synchronization in many-core neural network accelerators

This application describes a network-on-chip system that could be used in a hardware accelerator for accelerating neural network computations. An example NoC system may include a plurality of cores, and a plurality of data links connecting adjacent cores of the plurality of cores for transmitting data. The example NoC may further include a global synchronization switch connected to each of the plurality of cores. The global synchronization switch is configured to dynamically connect any pair of cores in the plurality of cores for transmitting control signals between the pair of cores.
Owner:MOFFETT TECH CO LTD

Neural network accelerator, hybrid convolution-vector operation processing system and computer implementation method

Neural network accelerators, hybrid convolution-vector operation processing systems, and computer-implemented methods are described. The neural network accelerator includes an instruction decoder configured to decode a neural network computing instruction from a processor into a weight load control signal, an activation load control signal, and a computing control signal; a plurality of weight selectors configured to obtain a weight according to a weight load control signal indicating whether the weight is obtained from the weight cache or the weight generator; a plurality of activation input interfaces configured to obtain an activation or a vector from the memory in accordance with an activation load control signal indicating whether to obtain the activation or the vector; and a plurality of circuit channels. The instruction decoder is further configured to, in response to the weights having a mode, instruct the plurality of weight selectors to obtain weights from the weight generator, rather than from the weight cache, to reduce memory access.
Owner:MOZI INT CO LTD

USB network accelerator and data transmission method and device

The invention provides a USB network accelerator and a data transmission method and device, and belongs to the technical field of computers and data transmission. Comprising a hardware processing unit, a software processing unit and a USB controller, the hardware processing unit is configured to perform packing operation or unpacking operation on a data packet which needs to be subjected to USB transmission between the network protocol stacks of the first equipment and the second equipment; the software processing unit is configured to execute a transmission control task of the data packet so as to send a data transmission command to the USB controller; and the USB controller is configured to transmit the data packet processed by the hardware processing unit between the first equipment and the second equipment based on the data transmission command. Therefore, the USB network accelerator accelerates processing of the USB network data in a software and hardware combination mode, the USB data transmission efficiency can be remarkably improved, and meanwhile resource consumption is reduced.
Owner:BEIJING X RING TECHNOLOGY CO LTD

Network-on-chip (NOC) for multi-core neural network accelerator

A network-on-chip system on a hardware accelerator for accelerating neural network computing is described. An example NoC system in an NN accelerator may include interconnected routers having routing control circuitry and cores coupled to the routers, respectively. The kernels are arranged in a matrix. Each row of cores is connected with the first one-way annular data link, and the directions of every two adjacent data links are opposite. Each column of cores is connected with the second one-way annular data link, and the directions of every two adjacent data links are opposite. In a given router of the plurality of routers, the route control circuitry is configured to: receive a data packet; converting the physical addresses of the given router and the target router into logic addresses; determining a routing port of the given router based on the logical address; and outputting the data packet through the routing port.
Owner:MOZI INT CO LTD

Lightweight neural network accelerator and quantization method

The application discloses a kind of lightweight neural network accelerator and quantification method, lightweight neural network accelerator, comprising: instruction scheduling module, convolution calculation module, wherein;The instruction scheduling module is connected with the convolution calculation module, for sending operation instruction to the convolution calculation module and executing convolution calculation to input data;The convolution calculation module is used to receive operation instruction, and according to operation instruction, convolution calculation is executed to the input data, and convolution inference result is output;Wherein, the convolution calculation module includes shifter, each convolution layer of the convolution calculation module executes convolution calculation to the input data, and 2 power n shift is carried out through the shifter, n is natural number greater than or equal to 0.The model calculation architecture of accelerator can be ensured to be efficient, low power consumption, high precision by the application.
Owner:SMARTSENS TECH (SHANGHAI) CO LTD

Training acceleration system, training acceleration method, and electronic device

This application discloses a training acceleration system, training acceleration method, and electronic device, particularly relating to the field of accelerated computing technology. It includes an external storage module and an acceleration module. The acceleration module comprises a processing module and a control module. The control module includes a storage protocol controller, an on-chip storage module, and a neural network accelerator. The external storage module stores training-related data. The processing module sends training task parameters for the neural network to the control module. The storage protocol controller generates read instructions for the external storage module based on the training task parameters and reads target data from the training-related data according to the logical block address specified in the read instructions. The on-chip storage module stores the target data. The neural network accelerator reads the target data and trains the neural network based on the target data. This solution addresses the problem of high data transmission latency and low training efficiency in existing technologies, achieving the technical effect of reducing data transmission latency and improving training efficiency.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD