Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

49 results about "Web accelerator" patented technology

A web accelerator is a proxy server that reduces web site access time. They can be a self-contained hardware appliance or installable software. Web accelerators may be installed on the client computer or mobile device, on ISP servers, on the server computer/network, or a combination. Accelerating delivery through compression requires some type of host-based server to collect, compress and then deliver content to a client computer.

Hybrid network-on-chip (NOC) for thread synchronization in many-core neural network accelerators

This application describes a network-on-chip system that could be used in a hardware accelerator for accelerating neural network computations. An example NoC system may include a plurality of cores, and a plurality of data links connecting adjacent cores of the plurality of cores for transmitting data. The example NoC may further include a global synchronization switch connected to each of the plurality of cores. The global synchronization switch is configured to dynamically connect any pair of cores in the plurality of cores for transmitting control signals between the pair of cores.
Owner:MOFFETT TECH CO LTD

MAC unit supporting input activation and weight double-end sparsity

The invention discloses an MAC unit supporting input activation and weight double-end sparsity, and relates to the technical field of neural network accelerator hardware design. Aiming at the technical defects existing in a sparse matrix acceleration scheme, the scheme is adopted, an activation value and non-zero elements of a weight matrix are stored through a special compressed data register, corresponding zero value distribution is recorded through a matched sparse bitmap register, and a joint sparse bitmap marking a double-non-zero effective calculation position is generated through dual-port joint operation; and the control and address generation subunit scans the joint sparse bitmap and outputs an address and an enable signal, the MAC calculation subunit is driven to read non-zero data to complete multiply-accumulate operation, and finally, an operation result is compressed and coded through an output end sparse encoder, so that efficient transmission of interlayer sparse data streams is realized. According to the method, input activation and weight double-end sparsity can be utilized at the same time, end-to-end sparse data flow is supported, zero value calculation can be dynamically skipped, and the energy efficiency ratio and the calculation throughput of neural network reasoning are improved.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

System and architecture of pure functional neural network accelerator

An accelerator circuit includes a control interface to receive a stream of instructions, a first memory to store an input data, and an engine circuit. The engine circuit includes a dispatch circuit to decode an instruction of the stream of instructions into a plurality of commands and a plurality of queue circuits. Each of the plurality of queue circuits supports a queue data structure to store a respective one of the plurality of commands decoded from the instruction, and a plurality of command execution circuits. Each of the plurality of command execution circuits is to receive and execute a command extracted from a corresponding one of the plurality of queues.
Owner:HUAXIA GENERAL PROCESSOR TECH INC

HARDWARE TIMER MANAGER THAT SUPPORTS MULTIPLE ACCURACY LEVELS

A hardware timer manager can be used in a network accelerator. The timer manager can support multiple timer precision levels. A timer manager can divide a total group of supported timer entries into multiple precision groups in a memory. The timer manager can check a programmed number of timer entries in a highest precision group before checking a single programmed timer entry in the next group, and repeat this check until a timer entry in the lowest precision group is checked, and then repeat from the beginning.
Owner:TEXAS INSTRUMENTS INC

Efficient data movement for ai accelerators

Efficient data movement in neural network accelerators operating within virtualized memory systems is challenged by high address translation latency and unique data access patterns. To address this challenge, address translation prefetch (ATP) mechanisms can be implemented to proactively translate virtual memory addresses before data movement. ATP can be performed in advance of any data movement or concurrently with data movement while being throttled by page transition in the data movement request stream. The ATP mechanism can enforce quotas on outstanding ATP requests, independently for read and write streams, to preserve resources for other processes running on the neural network accelerator. In dealing with competing ATP requests, the mechanism can employ weighted arbitration to balance between different types of ATP requests, utilizing a programmable ratio. The ATP mechanisms enable scalable, high-throughput neural network inference in virtualized environments, addressing data movement bottlenecks in neural network accelerator deployments.
Owner:INTEL CORP

Hardware architecture, computing method and device for 3D point cloud neural network algorithm

The application relates to the technical field of point cloud neural network, in particular to a hardware architecture, a calculation method and equipment for a 3D point cloud neural network algorithm, wherein the hardware architecture comprises: off-chip storage; a mapping module, the mapping module is provided with distance filtering technology and / or output priority mapping calculation technology to reduce the off-chip memory access amount of the off-chip storage when calculating the mapping operation of the 3D point cloud neural network algorithm, and a mapping relationship between input and output is generated according to the mapping operation; and a calculation module, the calculation module is provided with an elastic array architecture, the elastic array architecture is adjusted according to different scales of calculation tasks in the 3D point cloud neural network algorithm, and the weight and the corresponding input feature are taken out according to the mapping relationship to perform matrix operation to obtain the output feature. Thus, the problems in the prior art that the off-chip memory access amount of the point cloud neural network accelerator is large, the utilization rate of the calculation unit is low, the processing speed of the accelerator is slow, the expansibility and flexibility are poor, and the actual needs cannot be met are solved.
Owner:TSINGHUA UNIVERSITY

Low-computational complexity neural network accelerator based on superposed pilot frequency

The invention relates to the technical field of wireless communication, in particular to a hardware accelerator of a channel estimation model based on a neural network in the field of wireless communication, the accelerator adopts a software and hardware collaborative design, and the core architecture comprises a control processing unit, a computing unit, an on-chip memory, an off-chip memory and an AXI interface, the control processing unit is integrated with a least square method module, a control module and a hardware interface module, the control module adopts a ping-pong buffer mechanism to schedule tasks, the computing unit comprises an OFDM decoder, a fixed-point quantization module, a superposition pilot frequency module, a channel judgment module, a neural network module and an output module, and by utilizing a superposition pilot frequency technology, the channel judgment module and the neural network module are integrated into the control processing unit. Through low-bit wide fixed-point quantization, sparse weight decoding and a special pipeline computing architecture, a computing link is constructed in a neural network module, and the scheme is used for solving the problem of high computing complexity when a deep learning channel estimation method is deployed in a resource-limited edge scene in the prior art.
Owner:CHANGCHUN UNIV OF SCI & TECH

Lookup table engine supporting multiple protocols

PCT designated stageWO2025179093A1TransmissionLookup tableByte
A lookup table engine (100) may be implemented in hardware logic and yet provide operation for a multitude of different communication protocols and packet types. A lookup table memory (112) may be populated with rules that indicate, among other things, which bytes of a particular packet to extract, which comparisons to make to those extracted bytes, and actions to be taken based on the results of the comparisons. The lookup table engine may be implemented within a network accelerator (800), which receives a packet, and a processor core (808) of the network accelerator may offload lookup operations to the lookup table engine.
Owner:TEXAS INSTRUMENTS INC

Training neural network accelerators using mixed precision data formats

Techniques related to training neural network accelerators using mixed precision data formats are disclosed. In one example of the disclosed technology, a neural network accelerator is configured to accelerate a given layer of a multi-layer neural network. The input tensor of a given layer can be converted from a common precision floating point format to a quantized precision floating point format. A tensor operation may be performed using the converted input tensor. The tensor operation result can be converted from a block floating point format to a common precision floating point format. The converted result can be used to generate an output tensor of the layer of the neural network, wherein the output tensor is in a normal precision floating point format.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

An FPGA-based target detection neural network accelerator and a target detection system

The application discloses an FPGA-based target detection neural network accelerator and a target detection system, the accelerator is arranged on an FPGA and comprises a control module, an input buffer module, a parameter buffer module, a systolic array module, a post-processing module and a pooling module; the control module comprises an instruction memory and a controller; the instruction memory is used for storing an instruction set of a neural network; the controller is used for generating control instructions according to the instruction set and controlling the operation of each module; the input buffer module is used for storing input feature data blocks of the target detection neural network; the parameter buffer module is used for storing bias data and weight data of the target detection neural network; the systolic array module is used for performing convolution operation; the post-processing module is used for processing the convolution calculation result; the pooling module is used for performing maximum pooling operation on data; and the output buffer module is used for reading and writing cache of data. The application improves the inference efficiency of the accelerator by configuring parameters of the target detection neural network, and realizes the high-performance and high-flexibility target detection neural network accelerator.
Owner:GUANGDONG UNIV OF TECH

Neural network accelerators, neural network acceleration methods, and devices

ActiveCN117273094BData segmentTerm memory
This invention relates to the field of artificial intelligence chips, providing a neural network accelerator and a method and apparatus for accelerating neural networks. The neural network accelerator includes: a direct memory access controller that continuously reads target data stored in external memory according to data arrangement order and sends it to a data warping and control module; the data warping and control module that generates at least one target data segment based on the received data according to its bit width, the bit width of the computing unit, and the convolution parameters; an accelerated computing module that performs accelerated computing on the target data segment from the data warping and control module based on at least some computing units, and outputs the computing results of the computing units to a data rearrangement module; and a data rearrangement module that merges the computing results of each computing unit according to the data arrangement order required for the accelerated computing of the next layer of the neural network as the input for the accelerated computing of the next layer. This improves acceleration performance.
Owner:INST OF SEMICONDUCTORS - CHINESE ACAD OF SCI

Hardware Queue Manager Serving Multiple Processor Cores

A hardware-based queue manager is implemented in a network accelerator. The queue manager may receive input from a processor core into a set of registers designated for that processor core. Other sets of registers may be designated for other processor cores. The queue manager reads the input in the set of registers, which triggers the queue manager to begin a transaction, such as reading or writing to a queue data structure in a data memory. The queue data structure may handle multiple sizes of queue data structures.
Owner:TEXAS INSTRUMENTS INC

Reconfigurable hardware buffers in neural network accelerator framework

Embodiments of the present disclosure relate to reconfigurable hardware buffers in a neural network accelerator framework. A convolution accelerator framework (CAF) has a plurality of processing circuits including one or more convolution accelerators, a reconfigurable hardware buffer configurable to store data of a variable number of input data lanes, and a flow switch coupled to the plurality of processing circuits. The reconfigurable hardware buffer has a memory and control circuitry. The number of the variable number of input data lanes is associated with an execution epoch. During processing of the execution epoch, the flow switch streams data of the variable number of input data lanes between a processing circuit of the plurality of processing circuits and the reconfigurable hardware buffer. The control circuitry of the reconfigurable hardware buffer configures the memory to store the data of the variable number of input data lanes, the configuration including allocating a portion of the memory to each of the variable number of input data lanes.
Owner:STMICROELECTRONICS SRL

Binary network accelerator with floating point coprocessor and programmable control method

According to the binary network accelerator with the floating-point coprocessor and the programmable control method, the floating-point coprocessor is integrated on a chip, so that the supporting capacity for extra floating-point operation is remarkably improved. The accelerator adopts a three-stage pipeline architecture, floating point operation before convolution, binary convolution and floating point operation after convolution are respectively processed, and all stages are coordinated by a special control module, so that efficient parallel processing is realized. The floating point coprocessor has time and space parallelism, is composed of a plurality of vector group units, and supports general floating point operation. Accelerator programmable control can prevent circuit layout and wiring from being carried out again when working loads are switched. The accelerator is divided and executes two types of sub-graphs in a pipeline mode, various network topologies and residual connection are compatible, the limitation that a traditional accelerator is fixed in structure, poor in flexibility and the like is overcome, and the method can be widely applied to high-precision visual tasks.
Owner:XIDIAN UNIV

Filtering method, device and equipment of convolution-based FIR (Finite Impulse Response) digital filter and medium

The invention relates to the technical field of FIR digital filters, and discloses a filtering method, device and equipment of a convolution-based FIR digital filter and a medium, and the method comprises the steps: obtaining a one-dimensional convolution kernel and activation data corresponding to an N-order FIR digital filter; through LineBuf, filling operation is carried out on the activated data to overcome the boundary effect, and filled sub-data is obtained; and performing convolution operation based on the convolution kernel and the subdata through the MAC array to obtain a filtering result. The method is based on the convolution method and designs the LineBuf module loaded in stages, and benefits from the advantage, and a large amount of resource overhead can be reduced under the degree of parallelism. Besides, benefited from the advantages of the convolution method, the calculation parallelism degree is improved under the condition that the register overhead is not greatly increased, delay caused by software processing data filling is avoided through the LineBuf module, and high-speed calculation can be carried out under the condition of low resources. And meanwhile, the neural network accelerator and a common neural network accelerator are fused and share one MAC array, so that the hardware overhead is further reduced.
Owner:SHENZHEN UNIVERSITY OF ADVANCED TECHNOLOGY

Method for improving network security, electronic equipment and readable medium

The embodiment of the invention provides a system for improving network security, and the system comprises a driver of a first space of an operating system, and the driver is configured to analyze a network flow packet to obtain a first network flow packet, so that the operating system can read the content in the first network flow packet; the network accelerator of the first space of the operating system comprises a data layer and a man-in-the-middle attack tool; the man-in-the-middle attack tool is configured to filter malicious packets in the first network traffic packet according to the attribute of the first network traffic packet to obtain a second network traffic packet; the TCP / IP model of the first space of the operating system is configured to analyze the second network flow packet so as to establish communication with a first application layer of a second space of the operating system; the first application layer of the second space of the operating system comprises a conversion module, and the conversion module is configured to perform data conversion on the second network flow packet so as to enable the first application layer to run.
Owner:SIEMENS AG

Electric power metering terminal

The invention discloses an electric power metering terminal, and relates to the technical field of intelligent power grids. The system comprises an intelligent electric meter terminal, an edge AI gateway and a cloud platform. The intelligent electric meter terminal performs high-frequency sampling; the edge AI gateway realizes anomaly detection and load decomposition through an STM32H7 processor, a neural network accelerator, a lightweight ST-CNN model and an improved Seq2Point model, and communicates with the cloud platform; and the cloud platform realizes data aggregation and management. The method comprises the steps of anomaly detection, load decomposition and energy efficiency optimization. According to the invention, the problems of insufficient real-time performance, low identification precision and weak edge capability in the prior art are solved, millisecond-level response, high-precision load decomposition and low-power-consumption edge calculation are realized, and the annual average electric charge of a family can be reduced by 15-18%.
Owner:XIAN LIANGLI INSTR & METER

Lookup Table Engine Supporting Multiple Protocols

A lookup table engine may be implemented in hardware logic and yet provide operation for a multitude of different communication protocols and packet types. A lookup table memory may be populated with rules that indicate, among other things, which bytes of a particular packet to extract, which comparisons to make to those extracted bytes, and actions to be taken based on the results of the comparisons. The lookup table engine may be implemented within a network accelerator, which receives a packet, and a processor core of the network accelerator may offload lookup operations to the lookup table engine.
Owner:TEXAS INSTRUMENTS INC

Page fault support for virtual machine network accelerators

ActiveUS12712829B2Data packEngineering
Systems and methods for supporting page faults for virtual machine network accelerators. In one implementation, a processing device may receive, at a network accelerator device of a computer system, from a network, a first incoming packet and a second incoming packet. Responsive to receiving a first notification that an attempt to store the first incoming packet at a first buffer of a plurality of buffers associated with the network accelerator device caused a page fault, the processing device may store the first incoming packet at a second buffer and append a first identifier of the first buffer to a faulty buffer data structure. Responsive to receiving a second notification indicating a resolution of the page fault, the processing device may remove the first identifier from the faulty buffer data structure. The processing device may store the second incoming packet at the first buffer. The processing device may forward, to a driver of the network accelerator device, a second identifier of the second buffer and the first identifier of the first buffer.
Owner:RED HAT INC

A BERT network accelerator based on FPGA bit serial systolic array

This invention discloses a BERT network accelerator based on an FPGA-based bit-serial systolic array, comprising: an input data controller, a systolic array operation block, an intermediate data controller, an intermediate data buffer, an output encoder, and an output controller. This accelerator uses weighted data that has undergone bit-consistent quantization and compression for computation. The systolic array operation block employs bit-serial PE units, flexibly adapting to mixed-precision quantization and matrix multiplication calculations with low effective bit counts. The intermediate data controller is compatible with computations with and without BIAS and drives the intermediate data buffer. The output encoder performs bit-consistent quantization and compression on-chip on the data from the intermediate data buffer before sending it to the output controller. This accelerator fully utilizes the bit sparsity in the weighted data for bit-serial multiplication calculations and on-chip data compression, reducing logic and storage resource consumption; its compatibility with multiple computation modes improves the versatility of the BERT network FPGA accelerator and reduces computational costs.
Owner:BEIJING UNIV OF TECH

Resource management method and heterogeneous chip

The invention relates to a resource management method, a heterogeneous chip and computer equipment. The method is applied to a bridging unit connected between a processor unit and a neural network accelerator, and comprises the following steps: acquiring physical state data and calculation state data of the neural network accelerator; according to the task queue data, priority arbitration is carried out on tasks of a to-be-executed task queue of the neural network accelerator, an arbitration result is obtained, and the arbitration result comprises existence of high-priority hard real-time tasks or absence of the high-priority hard real-time tasks; calculating state data and an arbitration result according to the physical state data, and generating a voltage regulation signal and a frequency regulation signal; and adjusting the voltage, the frequency and the clock switching state of the neural network accelerator according to the voltage adjusting signal and the frequency adjusting signal. By adopting the method, the energy efficiency, the real-time performance and the management uniformity of the heterogeneous computing system can be improved.
Owner:CCORE TECH CO LTD

Neural network accelerator, data processing device and neural network acceleration method

ActiveCN113780541BPhysical realisationProcessing elementWeb accelerator
The present application provides a neural network accelerator, comprising a first processing unit and a second processing unit; a first data transfer module in the first processing unit is connected to a second data transfer module in the second processing unit; the first processing unit sends first target data to the second data transfer module via the first data transfer module, and / or the first processing unit receives second target data sent by the second data transfer module via the first data transfer module. The present application also provides a data processing device and a neural network acceleration method.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Hardware queue manager serving multiple processor cores

PendingCN122663559AEngineeringData memory
A hardware-based queue manager (100) is implemented in a network accelerator (500). The queue manager can receive input from a processor core into a register set designated for the processor core. Other register sets can be designated for other processor cores. The queue manager reads the input in the register set, which triggers the queue manager to start a transaction, such as reading or writing to a queue data structure in a data memory (114). The queue data structure can handle multiple sizes of queue data structures.
Owner:TEXAS INSTRUMENTS INC

Accelerator device with automatic precision switching based on the sparsity rate

Device for automatically adjusting the arithmetic precision in a neural network accelerator, comprising: a) at least one computing unit (2) with a plurality of processing elements (1) supporting at least two precision modes (INT4, INT8, INT16); b) a sparsity measurement circuit in each computing unit which continuously measures the rate of skipped zero operations (4) in relation to the total number of operations (3) during the execution of multiply-accumulate operations within a configurable measurement window; c) a precision controller in each computing unit which automatically switches the arithmetic precision between the precision modes based on the measured sparsity rate (5);andd) a hysteresis logic in the precision controller with separate thresholds for switching up to lower precision (8, 10) and switching down to higher precision (9, 11), where TH_UP > TH_DOWN, to prevent oscillation at borderline sparsity rate.;
Owner:NEUROVEXON UG (HAFTUNGSBESCHRÄNKT)

Processing method and processing system utilizing a multi-core neural network accelerator

The invention is a processing method for carrying out a first processing task flow and at least one further processing task flow on a multi-core neural network accelerator having processing units (11, 12), a command scheduler, and a memory, wherein each processing task flow comprises commands (20, 21, 22, 23) and dependencies between at least some of the commands (20, 21, 22, 23). Each command (20, 21, 22, 23) is executed by one particular processing unit (11, 12). The scheduling of the commands (20, 21, 22, 23) is determined by a merged command schedule, maintaining the dependencies within the respective processing task flows and reducing idle times of the processing units (11, 12). The merged command schedule is determined by a command schedule merger unit based on the input command schedules and their timelines. The invention also relates to a processing system for carrying out the method.
Owner:AIMOTIVE KFT

Neural network accelerators performing operations with mixed format weights

The invention relates to a neural network accelerator performing operations with mixed format weights. The data processing unit may include a memory, a processing element (PE), and a control unit. The memory may store a weight block in a weight tensor of a neural network operation. Each weight block has an input channel (IC) dimension and an output channel (OC) dimension, and includes sub-blocks. The sub-block includes one or more weights having a first data precision and one or more other weights having a second data precision. The second data precision is lower than the first data precision. The control unit may distribute different sub-blocks to different PEs. The PE may receive the sub-block and perform a first MAC operation on a weight having a first data precision and a second MAC operation on a weight having a second data precision. The first MAC operation may consume more computation cycles or more multipliers than the second MAC operation.
Owner:ALTERA CORP

Cyclic Redundant Spare Testing (CREST)

PendingUS20260188414A1MultiplexingMultiplexer
A fault-tolerant computing architecture performs cyclic runtime testing and repair in column-oriented processing arrays. Operational columns are periodically paired with test columns that receive identical inputs, and their outputs are compared to detect faults while computation proceeds at full throughput. Upon detecting a validated mismatch, control circuitry substitutes a spare column for the faulty column. Inter-column multiplexers enable defect-location at a segment granularity by redirecting partial outputs to subsequent columns during diagnostic steps, after which normal routing is restored with the spare in place. The approach supports threshold-based validation to filter transient errors, pre-validation of spares, and scheduling at layer boundaries to avoid latency impact. The system enhances reliability of large neural-network accelerators without downtime and enables graceful degradation through dynamic substitution within the array.
Owner:SILVEBROOK KIA

Large language model operator conversion system and method and AI accelerator

The invention provides a large language model operator conversion system and method and an AI accelerator, the system comprises an operator identification module, a conversion module and a model reconstruction module, the operator identification module analyzes a computational graph structure of an input model to identify convertible operators meeting preset conversion conditions in the computational graph structure; the conversion module converts the convertible operator into an equivalent operator sequence based on a preset conversion strategy; and the model reconstruction module replaces nodes corresponding to convertible operators in the computational graph of the input model with an equivalent operator sequence, updates the connection relation and the weight initializer, and generates an updated model. According to the system, pipeline processing of model optimization is realized through an identification rule and a conversion strategy, and adaptation for different hardware platforms can be realized by updating a preset conversion condition and a conversion strategy, so that the problems of low deployment efficiency and low expandability on a CNN (Convolutional Neural Network) accelerator are solved.
Owner:NANJING UNIV

Network architecture, corresponding vehicle and method

The bandwidth of SOC interfaces is exploited while minimizing the number of physical ports via a networking accelerator for use on board a vehicle, for instance, that comprises: media access control (MAC) controller circuitry configured to provide a MAC port layer to control exchange of information, wherein the exchange of information comprises data flow transmission to virtual machine ports (VMPs) over a data link; virtual machine transmission (VM Tx) bridge circuitry configured to handle transmission data flow to the VMPs; transmission router / switch circuitry configured to route / switch data flow from the MAC controller circuitry to the VM Tx bridge circuitry; and queue handler circuitry configured to provide queue management for data flow between the MAC controller circuitry and the VM Tx bridge circuitry. The VM Tx bridge circuitry comprises virtual destination address circuitry configured to implement router / switch virtualization in the transmission router / switch circuitry with a virtual machine transmission descriptor based on a combination of a virtual machine port (VMP) tag indicative of a physical resource in the queue handler circuitry selectable for data flow transmission, and a virtual machine extended identifier (VMEID).
Owner:STMICROELECTRONICS INT NV

Neural network activation compression with narrow block floating-point

Apparatus and methods for training a neural network accelerator using quantized precision data formats are disclosed, and in particular for storing activation values from a neural network in a compressed format for use during forward and backward propagation training of the neural network. In certain examples of the disclosed technology, a computing system includes processors, memory, and a compressor in communication with the memory. The computing system is configured to perform forward propagation for a layer of a neural network to produced first activation values in a first block floating-point format. In some examples, activation values generated by forward propagation are converted by the compressor to a second block floating-point format having a narrower numerical precision than the first block floating-point format. The compressed activation values are stored in the memory, where they can be retrieved for use during back propagation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC