Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

25 results about "Systolic array" patented technology

In parallel computer architectures, a systolic array is a homogeneous network of tightly coupled data processing units (DPUs) called cells or nodes. Each node or DPU independently computes a partial result as a function of the data received from its upstream neighbors, stores the result within itself and passes it downstream. Systolic arrays were invented by H. T. Kung and Charles Leiserson who described arrays for many dense linear algebra computations (matrix product, solving systems of linear equations, LU decomposition, etc.) for banded matrices. Early applications include computing greatest common divisors of integers and polynomials. They are sometimes classified as multiple-instruction single-data (MISD) architectures under Flynn's taxonomy, but this classification is questionable because a strong argument can be made to distinguish systolic arrays from any of Flynn's four categories: SISD, SIMD, MISD, MIMD, as discussed later in this article.

A lightweight neural network processor storage architecture co-optimization method

ActiveCN121809563BComputer hardwareNetwork processor
The present application relates to neural network inference hardware technical field, particularly to a kind of lightweight neural network processor storage architecture collaborative optimization method, comprising: step 1: the difference of on-chip data on bandwidth demand and access mode is analyzed, and the differential single-port storage setting is carried out to the buffer storage of each data in neural network processor;Step 2: the GEMM module of neural network processor is based on semi-pulsating array setting, and ALU operation fusion processing module is used, to reduce the storage bandwidth and bit width of data;Step 3: the storage structure is optimized by introducing the dynamic adjustment strategy based on timing analysis, to output the optimized neural network processor.The present application guarantees the computing performance, significantly improves the area efficiency and access efficiency of storage system, realizes the collaborative optimization of bandwidth, area and power consumption.
Owner:ZHEJIANG UNIV

Systems and methods for accelerating neural network convolution and training

ActiveCN114902242BCounter propagationNeural network nn
A specialized integrated circuit for artificial neural networks is integrated with high bandwidth memory. The neural network includes a systolic array of interconnected processing elements, including upstream processing elements and downstream processing elements. Each processing element includes a pair of input / output ports for concurrent forward and backward propagation. The processing elements can be used for convolution, in which case the pair of input / output ports can support fast and efficient scanning of a kernel over activations.
Owner:RAMBUS INC

Methods and apparatus for deep learning

A method and apparatus for deep learning are disclosed. The apparatus for deep learning includes a processor configured to support a plurality of different operation modes, the processor including a systolic array having a plurality of multiplier accumulator (MAC) units, and a control circuit configured to control a selection operation of the plurality of MAC units and a data movement between the plurality of MAC units, respectively, for each of the plurality of different operation modes.
Owner:SAMSUNG ELECTRONICS CO LTD

Processing unit, systolic array, data processing method, and electronic device

PCT designated stageWO2026144749A1Processing elementData transmission
The present disclosure provides a processing unit of a systolic array, the systolic array and a data processing method, and an electronic device. The processing unit includes a first set of data transmission paths for transmitting data in a first dimension and a second set of data transmission paths for transmitting data in a second dimension. The first set of data transmission paths are configured to transmit first data. The second set of data transmission paths is configured to transmit second data. The first set of data transmission paths includes a first data transmission path for transmitting data in a first direction and a second data transmission path for transmitting data in a second direction. The second set of data transmission paths includes a third data transmission path for transmitting data in a third direction and a fourth data transmission path for transmitting data in a fourth direction.
Owner:SMARTER SILICON (SHANGHAI) TECH CO LTD

Clocking a systolic array on both edges of a clock signal

The present disclosure is directed to an apparatus that includes a systolic array having an initial systolic stage that is clocked at a first edge of a clock signal. The apparatus further includes control logic configured to clock a first subset of a plurality of additional systolic stages of the systolic array at the first edge of the clock signal. The control logic is further configured to clock a second subset of the plurality of systolic stages at a second edge of the clock signal.
Owner:QUALCOMM INC

An end-side accelerator and design method of a distributed MoE large model architecture

This invention belongs to the field of data processing and discloses an edge accelerator and its design method for a distributed MoE large model architecture. The edge accelerator includes: a distributed vector systolic array module for acquiring initial text data and performing attention calculations to obtain initial text features; a nonlinear operation module for calculating expert scores on the initial text features to obtain the target expert score corresponding to each expert in the routing network; and a MoEGate module for performing expert selection and data rearrangement processing based on the target expert scores to obtain the mapping relationship between experts and tokens, and also for performing expert inference on the initial text features based on the mapping relationship between experts and tokens to output the text inference result. This invention, by adopting a distributed integrated architecture and integrating the MoEGate module, provides a sufficient bandwidth foundation and significantly improves inference throughput and energy efficiency, greatly accelerating the processing efficiency of text data.
Owner:SHENZHEN MAITEXIN TECH CO LTD

A neural network oriented accelerator, processor

The application provides a neural network-oriented accelerator, which comprises a main processing core, an input buffer, a controller, a main systolic array and an output buffer, the input buffer is used for storing feature maps and weight parameters of a neural network, the controller is used for generating operation signals to read the feature maps and the weight parameters from the input buffer, the main systolic array is used for receiving the feature maps and the weight parameters and performing calculation to obtain an output sequence; a sorter is used for reordering each element in the output sequence to select a to-be-detected sequence, a lockstep processing core is used for recalculating the value of each element in the to-be-detected sequence through a lockstep systolic array to obtain a check sequence; a check logic module is used for comparing whether the values of the same elements in the to-be-detected sequence and the check sequence are the same; and an error recovery module is used for recalculating the values of all elements in the to-be-detected sequence that do not pass the check according to a preset recalculation mechanism.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

A systolic array weight input control system

ActiveCN115455997BComputing operations for integral formationMedicineLogic cell
The present application relates to a kind of weight input control system of pulsating array, belong to pulse array control technical field.The weight input control system of pulsating array includes: control unit, input storage unit, weight storage unit, logic unit, pulsating array unit and output storage unit;The control unit is connected with the input storage unit, the logic unit and the pulsating array unit respectively;The logic unit is connected with the weight storage unit;The pulsating array unit is connected with the input storage unit, the weight storage unit and the output storage unit respectively;Based on this structure setting, the present application can reduce image data transmission time, and then improve computing efficiency.
Owner:NANJING INST OF INTELLIGENT TECH INST OF MICROELECTRONICS OF THE CHINESE ACAD OF

Neural network computation apparatus having systolic array

A neural network computation apparatus includes a first processing block including a plurality of processing units that each perform a matrix multiplication operation on input data and weights, and a second processing block including a plurality of element-wise operation processing groups. The element-wise operation processing group selectively perform a first neural network computation operation and a second neural network computation operation. The first neural network computation operation comprises the matrix multiplication operation on the input data and the weights and an activation operation on a result value of the matrix multiplication operation, and the second neural network computation operation comprises an activation operation on the result value of the matrix multiplication operation, which is transferred from the first processing block, and an element-wise operation.
Owner:SK HYNIX INC

Matrix multiplication conversion method, apparatus, medium, device and product

The application discloses a matrix multiplication conversion method and device based on systolic array convolution calculation, a medium, equipment and a program product, and belongs to the technical field of neural network processors. The method mainly comprises the following steps: according to the input channel number of an input feature map of a systolic array, cutting matrix row data of a first matrix, and inputting the cut matrix row data of the same row into the systolic array in the form of row data; according to the input channel number and the output channel number of the weight of the systolic array, cutting matrix column data of a second matrix, and inputting the cut matrix column data of the same row into the systolic array in the form of column data; and calculating the product of the first matrix and the second matrix by using the systolic array. The application improves the data reuse rate of matrix multiplication operation, reduces memory access, improves the cache utilization rate, and optimizes data locality.
Owner:CCORE TECH CO LTD

Structured sparse matrix acceleration in systolic arrays

Methods, systems, and apparatus, including computer-readable storage media for hardware-accelerated fine-grained sparse computation. The accelerator provides for improved performance for structured fine-grained sparse AI workloads, for example by accelerating sparse matrix multiplication required to execute or train AI models. Sparse data is compressed to remove zero-valued elements before being streamed into a matrix multiplication unit (MXU) of the accelerator. The accelerator stores a gains matrix, which can be the matrix for multiplying with the received input matrix. The accelerator uses an index array mapping locations of elements in the compressed matrix with locations of elements in the matrix's precompressed form, to generate a multiplier matrix from the gains matrix. Aspects of the disclosure also provide for accelerated gains matrix loading in a hardware accelerator or other type of processor. The accelerator can load the gains matrix more efficiently in a compressed form, and then un-compress the matrix once loaded.
Owner:GOOGLE LLC

Systolic array including fused multiply accumulate with efficient prenormalization and extended dynamic range

Systems and methods are provided to perform multiply-accumulate operations of at least one normalized number in a systolic array. The systolic array can obtain a first input and detect that the first input is denormal. Based on determining the first input is denormal, the systolic array can generate a first normalized number by normalizing the first input. Processing elements of the systolic array can include a multiplier and an adder. The multiplier can multiply the first normalized number by a second normal or normalized number to generate a multiplier product and the adder can add an input partial sum to the multiplier product to generate an addition result.
Owner:AMAZON TECH INC

Model inference task processing system and model inference task processing method

The application discloses a model inference task processing system and method, relates to the technical field of model inference, and comprises a plurality of pre-filling stage hardware accelerators and a plurality of decoding stage hardware accelerators; the pre-filling stage hardware accelerator comprises a systolic array, and the systolic array processes a matrix multiplication matrix calculation task in a model inference process; the decoding stage hardware accelerator comprises a plurality of multiply-add trees, and the multiply-add trees process a matrix multiplication vector calculation task in the model inference process; the scale of the systolic array is the same as the total scale of the plurality of multiply-add trees; the computing instructions of the pre-filling stage hardware accelerator and the decoding stage hardware accelerator are compatible; and the computing node is used for determining a target pre-filling stage hardware accelerator and a target decoding stage hardware accelerator based on a model inference task, so as to process the model inference task. The technical effect of balancing TTFT and TPOT and improving the model inference efficiency is achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

FPGA-based MTLA-Transformer hardware accelerator

PendingCN122154768AInference methodsPhysical realisationFeed forward networkParallel computing
The application discloses an MTLA-Transformer hardware accelerator based on FPGA and relates to the technical field of FPGA acceleration and deep learning inference optimization. The accelerator comprises a controller module, an attention calculation module and a feedforward network module. The controller module is used for completing data scheduling of on-chip cache and off-chip memory and KV cache update management. The attention calculation module comprises a reusable matrix multiplication and addition calculation array, a position coding submodule, a HyperNet calculation submodule, a KV cache management submodule, a Softmax submodule and a residual normalization submodule. Projection calculation, fractional calculation and output projection and other matrix multiplication and addition operation multiplexing are realized through a systolic array. Position coding rotation factor precalculation lookup table, HyperNet position correlation result offline precalculation and Softmax scaling coefficient fusion are adopted to reduce online calculation amount and memory access overhead. The feedforward network module completes two-layer linear transformation and activation operation and realizes residual connection normalization processing. The application can improve incremental inference throughput and reduce hardware resource overhead under the premise of ensuring inference correctness and is suitable for edge end low-power real-time inference scenarios.
Owner:SUN YAT SEN UNIV +1

Column Aligned Vertical Reduction Network for FP4 Weight and FP8 Activation Systolic Arrays

PendingUS20260187027A1PathPingParallel computing
A column aligned vertical reduction network for systolic arrays propagates partial sums generated by processing elements operating on FP4 weights and FP8 activations. Each processing element performs fused multiply add operations using a two cycle pipeline with one cycle throughput. Column aligned accumulation nodes receive partial sums and forward them downward through a vertical reduction path. Physical and logical alignment of processing elements and reduction nodes maintains timing uniformity across columns. The network supports saturating partial sums, bypass routing, and multi-tier implementations combining logic and memory.
Owner:SILVEBROOK KIA

Feature map processing method and device based on systolic array and storage medium

The application provides a feature map processing method and device based on a pulsation operation array and a computer readable storage medium. The feature map processing method comprises: obtaining an input feature map to be processed and a convolution kernel used for three-dimensional convolution operation on the input feature map; decomposing the convolution kernel into a 1*1-dimensional sub-convolution kernel along a width direction and a height direction of the convolution kernel; and performing parallel operation on feature values of corresponding position points in a width-height plane of the input feature map and weight values of the sub-convolution kernel along a channel direction of the input feature map and the convolution kernel by using the pulsation operation array to obtain a partial convolution result; and accumulating the partial convolution result to obtain a corresponding feature value of an output feature map. In this way, the feature map processing device converts convolution kernels of different sizes into a data segmentation, scheduling control and hardware implementation scheme of 1*1 convolution kernel calculation, and solves the compatibility problem of different convolution kernel calculation.
Owner:ZHEJIANG DAHUA TECH CO LTD

Column-Array Systolic Computation with Accumulation During Execution (CASCADE)

PendingUS20260195408A1Computational scienceClock rate
A column-oriented compute architecture for neural-network inference is disclosed. Processing elements are arranged in vertical columns, each column storing stationary weight values and accumulating partial sums locally during a matrix-multiplication operation. For each processing row, an activation value is broadcast concurrently to all columns, allowing each processing element to multiply the broadcast activation by its stationary weight and accumulate the result in mixed-precision arithmetic such as FP4 weights and FP8 activations with FP8 accumulation. No partial sums are exchanged between columns. The architecture reduces interconnect power, simplifies data movement, and enables higher clock rates relative to conventional systolic arrays. In some embodiments, a three-dimensional stack integrates compute and broadcast-distribution dies coupled by hybrid bonding, and spare columns are dynamically substituted for faulty columns while maintaining the row-wise broadcast semantics.
Owner:SILVEBROOK KIA

A serial signal matrix multiplication operation circuit for Kalman filtering

The application discloses a serial signal matrix multiplication operation circuit for Kalman filtering and belongs to the field of Kalman filtering algorithm circuit implementation. The serial signal matrix multiplication operation circuit comprises 2n-1 D flip-flops, one logic or unit, n multiplier units and n processing units; the output result of the multiplier unit 1 is kg1*h1, kg2*h2...kg(n-1)*h(n-1), kg(n)*h(n), which is exactly the diagonal element of the matrix multiplication result and is used for subsequent matrix (I-kg*h) operation; the n multiplier units successively output each element in the same row of the matrix multiplication result every clock cycle, thereby facilitating subsequent systolic array matrix multiplication operation processing. The serial signal matrix multiplication operation circuit consumes certain hardware circuit resources to continuously output the matrix multiplication result in a pipeline mode, and the sequence of the result output greatly facilitates the processing of the subsequent result of Kalman filtering.
Owner:HUAZHONG UNIV OF SCI & TECH

Signal detection apparatus and method for multiple-input multiple-output systems

The application provides a signal detection device and method for a multiple-input multiple-output system, which can improve the throughput rate of MIMO signal detection of a receiver, and reduce the hardware cost of signal detection through processing structure optimization. Through complete systolic array design and data stream design, the intermediate matrix MEM in calculation is avoided, the hardware consumption is reduced, and complete high-speed MIMO signal detection processing is realized. Through the full systolic array design of MMSE detection, optimization of the structure and hardware resources is realized to realize high-speed array processing.
Owner:WUHAN ZHONGKE JINGSHANG INFORMATION TECHNOLOGY CO LTD

Systolic array matrix-multiply accelerator with row tail accumulation

An accelerator is accessed. The accelerator includes a systolic array of tiles that include a plurality of rows and a plurality of columns. Each tile in the systolic array of tiles includes one or more multiply-add units. The accelerator includes a row tail associated with each row within the plurality of rows. A first matrix is multiplied by a second matrix in the systolic array. The multiplying is based on a plurality of partial products. A row within the plurality of rows of the systolic array produces one or more partial products within the plurality of partial products. A last tile within the row forwards the one or more partial products to an associated row tail. The one or more partial products that were forwarded are accumulated with one or more previous partial products associated with the row by the associated row tail.
Owner:AKEANA INC

Systolic array with output rounding across multiple data streams

Systems and methods are provided to round the numbers produced by a systolic array. A rounder can receive a number from the systolic array and identify a data stream associated with the number from a plurality of data streams. The rounder can identify a random number generator. The random number generator may be associated with a random number sequence and may generate a next random number in the random number sequence based on a state value representing a position within the random number sequence. The data stream may be associated with a respective state value representing a current position for the data stream. Based on the current position for the data stream, the rounder can initialize a state value of the random number generator. The rounder can perform a rounding operation using the initialized state value of the random number generator.
Owner:AMAZON TECH INC

Scalable and configurable clustered systolic array

ActiveUS12664121B2Systolic arraysElectric digital data processingComputer architectureLogical combination
A scalable and configurable clustered systolic array is described. An example of apparatus includes a cluster including multiple cores; and a cache memory coupled with the cluster, wherein each core includes multiple processing resources, a memory coupled with the plurality of processing resources, a systolic array coupled with the memory, and one or more interconnects with one or more other cores of the plurality of cores; and wherein the systolic arrays of the cores are configurable by the apparatus to form a logically combined systolic array for processing of an operation by a cooperative group of threads running on one or more of the plurality of cores in the cluster.
Owner:INTEL CORP