Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

232 results about "Systolic array" patented technology

In parallel computer architectures, a systolic array is a homogeneous network of tightly coupled data processing units (DPUs) called cells or nodes. Each node or DPU independently computes a partial result as a function of the data received from its upstream neighbors, stores the result within itself and passes it downstream. Systolic arrays were invented by H. T. Kung and Charles Leiserson who described arrays for many dense linear algebra computations (matrix product, solving systems of linear equations, LU decomposition, etc.) for banded matrices. Early applications include computing greatest common divisors of integers and polynomials. They are sometimes classified as multiple-instruction single-data (MISD) architectures under Flynn's taxonomy, but this classification is questionable because a strong argument can be made to distinguish systolic arrays from any of Flynn's four categories: SISD, SIMD, MISD, MIMD, as discussed later in this article.

FPGA superposition processor acceleration system and method based on state space duality

The invention belongs to the field of machine learning, and discloses an FPGA (Field Programmable Gate Array) superposition processor acceleration system and method based on state space duality, which comprises a sparse predefined data acquirer, a reconfigurable systolic array, a partial sum cache, an element-by-element operation cache, a function calculation module, an on-chip memory management module and the like. Zero elements are eliminated through the sparse predefined data acquirer, redundant data are reduced, and redundant calculation is remarkably reduced; the reconfigurable systolic array flexibly supports multiple calculation modes, and the utilization rate of hardware resources is increased; the part and the cache realize cross-cycle accumulation and element-by-element operation cache integration results, the function calculation module completes nonlinear operation, and the on-chip memory management module optimizes result cache, so that on-chip operation of SSD calculation is ensured, off-chip memory access is remarkably reduced, and reasoning efficiency and energy efficiency are improved; by adopting the system, the memory occupation is effectively reduced, the element-by-element calculation efficiency is improved, and the sparse calculation redundancy is reduced, so that the reasoning process of the Mamba2 model is remarkably accelerated.
Owner:NINGBO ORIENTAL UNIV OF TECH (TEMPORARY NAME)

Flexible convolution operation accelerator based on systolic array

The invention provides a flexible convolution operation accelerator based on a systolic array. The flexible convolution operation accelerator is used for solving the technical problem that the hardware scale of an existing convolution operation accelerator is difficult to flexibly adjust according to different hardware platforms. The system comprises an input cache module, a pulsation matrix, an accumulation logic unit and an output cache module which are connected in sequence, the input cache module comprises a weight cache module and an image cache module, and the weight cache module and the image cache module are both connected with the pulsation matrix. The weight cache module, the image cache module, the pulsation matrix, the accumulation logic unit and the output cache module are all connected with the controller, and the controller and the output cache module are both connected with an external memory through a BUS. The systolic array and the flexible cache module are utilized, the functions of data storage and data format conversion can be achieved at the same time, hardware logic overhead is reduced, high memory access efficiency is achieved, and configuration of the accelerator scale and various convolution parameters is supported.
Owner:HENAN XUNGU TECH CO LTD +1

Systolic array with input reduction to multiple reduced inputs

Systems and methods are provided to perform multiply-accumulate operations of reduced precision numbers in a systolic array. Each row of the systolic array can receive reduced inputs from a respective reducer. The reducer can receive a particular input and generate multiple reduced inputs from the input. The reduced inputs can include reduced input data elements and / or a reduced weights. The systolic array may lack support for inputs with a first bit-length and the reducers may reduce the bit-length of a given input from the first bit-length to a second shorter bit-length and provide multiple reduced inputs with second shorter bit-length to the array. The systolic array may perform multiply-accumulate operations on each unique combination of the multiple reduced input data elements and the reduced weights to generate multiple partial outputs. The systolic array may sum the partial outputs to generate the output.
Owner:AMAZON TECH INC

Structured Sparse Matrix Acceleration In Systolic Arrays

Aspects of the disclosure are directed to hardware acceleration of structured sparse workloads with block quantization. A hardware accelerator can receive compressed input matrices, for example as part of a workload for training or processing a machine learning model. The hardware accelerator can multiply the compressed input matrix with a gains matrix loaded in one or more matrix multiply units (MXUs) of the hardware accelerator. The input matrices can be further provided in a block data type format, in which blocks of mantissas are represented with a single shared scaling factor. An MXU can multiply the block data, shift or cast the block data according to a shared scaling factor to generate an output product. To that end, block data type matrices exhibiting structured sparsity patterns can be accelerated without affecting the overall accuracy or quality of the output to the workload being processed.
Owner:GOOGLE LLC

Hardware accelerator with matrix block streaming

A hardware accelerator including tiles arranged in a systolic array. At each of the tiles, the systolic array receives a first input block that includes first input matrix elements of a first input matrix. In each of a plurality of multiplication iterations, at each of the tiles, the systolic array receives a respective second input block. The systolic array computes tile products of the first input matrix elements and second input matrix elements included in the second input blocks. The systolic array adds the tile products to column-wise partial sums and transmits the column-wise partial sums to subsequent tiles along accumulator rings included in array columns of the systolic array. In a subset of the multiplication iterations, the systolic array outputs product block rows of a product matrix. The product block rows each include product matrix blocks computed as rows of the column-wise partial sums.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Flexible array data loading

To improve utilization of a systolic array, each row of the array is provided with two or more general purpose row input buses. Each of the general purpose row input data buses can be operable to transfer either feature map (FMAP) input elements or weight values into the processing elements of the corresponding row of the array. The flexibility of the general purpose row input buses allows weights to be loaded in parallel into the array to speed up weight loading times to reduce the overall latency.
Owner:AMAZON TECH INC

Expert parallelism processing method and system of large language model based on MoE

The invention belongs to the field of machine learning, discloses an expert parallelism processing method and system for a large language model based on MoE, and realizes efficient parallel processing of the MoE model by dynamically distributing expert quantization bit width and sparse mode, predicting and prefetching to-be-activated expert parameters, grouping tokens to generate task queues and dynamically configuring hardware accelerators. Firstly, an importance score is calculated based on expert historical activation frequency, weight distribution and a model structure, so that quantization precision and a sparse proportion are adaptively allocated, and resource waste and precision loss of a unified strategy are avoided; secondly, predicting an expert to be activated by using a current layer hidden state and a historical activation sequence, loading parameters to a special cache in advance, reducing high-bandwidth memory access, and relieving bandwidth peak scrambling; moreover, tokens are grouped through a token-expert mapping table, a task queue is created, and dynamic configuration of a systolic array is combined, so that the problem of expert heterogeneity after compression is solved, and efficient and parallel hybrid precision matrix operation is ensured.
Owner:NINGBO ORIENTAL UNIVERSITY OF TECHNOLOGY +2

Systems and Methods for a Near Memory-Based Matrix Computation

Systems or methods of the present disclosure may provide an integrated circuit system that includes a programmable logic device that includes a clock, one or more local controllers, programmable logic units implementing a systolic array to compute a matrix multiplication, and embedded memory blocks. The embedded memory blocks include a single port random access memory (SPRAM). The one or more local controllers are configured to, on a first set of alternating clock cycles of the clock, load matrix sub-elements from two rows of a matrix into corresponding matrix element of the SPRAM. The one or more local controllers are configured to, on a second set of alternating clock cycles of the clock, read out the matrix elements from the SPRAM to the systolic array to compute the matrix multiplication.
Owner:ALTERA CORP

Processor architecture supporting in-memory matrix operation based on static memory and processor

The embodiment of the invention provides a processor architecture supporting in-memory matrix operation based on a static memory and a processor. Wherein the processor is provided with the provided processor architecture. A tensor core module in the processor architecture comprises a plurality of matrix operation units, a plurality of in-memory computing macro units are arranged in the matrix operation units, the plurality of in-memory computing macro units are connected by adopting a pulsation data path to form a two-dimensional pulsation array, and each in-memory computing macro unit comprises a plurality of parallel memory banks. And the memory bank comprises a static memory array supporting in-memory matrix operation. Therefore, according to the processor architecture provided by the scheme, the in-memory calculation macro unit based on the static memory is introduced to serve as the matrix operation unit to support in-memory matrix operation, and by adopting the matrix operation unit, the overall performance of the processor can be improved, and the high energy efficiency requirements of generative model reasoning and training can be better met.
Owner:PEKING UNIV

Systolic-CNN: an OpenCL-defined scalable runtime-flexible programmable accelerator architecture for accelerating convolutional neural network inference in cloud / edge computing

An OpenCL-defined scalable runtime-flexible programmable accelerator architecture for accelerating convolutional neural network (CNN) inference in cloud / edge computing is provided, referred to herein as Systolic-CNN. Existing OpenCL-defined programmable accelerators (e.g., field-programmable gate array (FPGA)-based accelerators) for CNN inference are insufficient due to limited flexibility for supporting multiple CNN models at runtime and poor scalability resulting in underutilized accelerator resources and limited computational parallelism. Systolic-CNN adopts a highly pipelined and paralleled one-dimensional (1-D) systolic array architecture, which efficiently explores both spatial and temporal parallelism for accelerating CNN inference on programmable accelerators (e.g., FPGAs). Systolic-CNN is highly scalable and parameterized, and can be easily adapted by users to efficiently utilize the coarse-grained computation resources for a given programmable accelerator. In addition, Systolic-CNN is runtime-flexible and can be time-shared to accelerate a variety of CNN models at runtime without the need to recompile the programmable accelerator kernel hardware or reprogram the programmable accelerator.
Owner:THE ARIZONA BOARD OF REGENTS ON BEHALF OF THE UNIV OF ARIZONA

Data processing method, system and terminal of multi-head potential attention model based on TTD compression

The invention discloses a data processing method and system for a multi-head potential attention model based on TTD compression and a terminal, and the method comprises the steps: constructing a large language model, and processing a plurality of linear layers in the large language model in a TTD compression and decomposition mode, thereby reducing the weight number in the model, and improving the data processing efficiency. And on a hardware level, targeted optimization is carried out on a data stream subjected to linear layer reasoning after TTD compression, so that a final model reasoning result is output. According to the method, a multi-head potential attention structure in a large language model is converted, so that the requirement for KV cache during model reasoning is reduced, the weight number is reduced, the long text output capacity of the model on edge equipment is improved, matrix calculation in the reasoning process is carried out subsequently by adopting a calculation structure of a group vector systolic array, and the reasoning efficiency is improved. And limited hardware resources are efficiently utilized.
Owner:SHENZHEN MAITEXIN TECH CO LTD

Efficient matrix engine architecture based on RISC-V matrix extension and calculation method

The invention provides a high-efficiency matrix engine (RVME) architecture based on RISC-V matrix extension and a calculation method, and the architecture comprises an instruction buffering and decoding module, a matrix loading / storage module, a matrix register file, a parallel outer product array and an element-by-element operation module; the matrix register file comprises a Tile register and an Acculator register; the storage modules are respectively used for storing an input matrix and an accumulation result and supporting efficient data access and parallel computing; the matrix loading / storage module significantly improves the data loading efficiency through cache line alignment and matrix transposition optimization; the instruction buffering and decoding module cooperates with a main processor through a reordering buffer area and an instruction buffer area to ensure efficient scheduling and execution of instructions. The parallel outer product array is adopted to replace a traditional systolic array, the idle period in the calculation process is eliminated through multicast data flow scheduling and a ping-pong buffer read-write mechanism, and matrix multiplication and addition operation with high calculation utilization rate and low delay is achieved.
Owner:SHANGHAI JIAOTONG UNIV

Multi-modal systolic array for matrix multiplication

A system and method for matrix multiplication using a systolic array configurable between a plurality of operating modes. A pulsation processor may receive a data type indicator for matrix multiplication. For a first data type, the systolic processor may load right-hand side data from a right-hand matrix register into data processing units of the systolic array between rows 0 and M-1, and pass respective rows of left-hand side data through corresponding rows of the systolic array between rows 0 and M-1. For the second data type, the pulsation processor may segment each element of the left-hand-side data and the right-hand-side data into a respective first element half and a respective second element half, and move each element half through a corresponding row of the pulsation array between rows 0 and 2M-1.
Owner:GOOGLE LLC

Data processing method and device for matrix calculation, pulsation matrix, electronic equipment and medium

The invention relates to a data processing method for matrix calculation, a pulsation matrix, a device, equipment and a medium. The data processing method comprises the steps that data in a first input matrix and data in a second input matrix are read, the first input matrix is provided with M rows and K columns, and the second input matrix is provided with K rows and N columns; inputting the read data of the first input matrix into the systolic matrix, and enabling the data of the first input matrix to be transversely and circularly propagated in the systolic array; inputting the read data of the second input matrix into the systolic array, and enabling the data of the second input matrix to be circularly propagated in the systolic array along the diagonal direction; the processing unit is used for executing multiply-accumulate operation on data flowing through the first input matrix and data flowing through the second input matrix, and a calculated intermediate result is resided in the processing unit; and in response to determining that the multiply-accumulate operation is completed, unidirectionally propagating the final calculation result residing in each processing unit along the diagonal direction so as to move the final calculation result out of the systolic array.
Owner:BEIJING WEIFAN INTELLIGENT TECHNOLOGY CO LTD

Self-attention mechanism calculation method, array, device and system and storage medium

The invention discloses a self-attention mechanism calculation method, a calculation array, a calculation device, a calculation system and a computer readable storage medium. The self-attention mechanism calculation method is applied to a self-attention mechanism calculation array, the self-attention mechanism calculation array comprises a plurality of product accumulators, the plurality of product accumulators form a systolic array, and the self-attention mechanism calculation method comprises the following steps: performing calculation through the plurality of product accumulators according to a feature tensor and a weight tensor to obtain an intermediate tensor; and a result tensor is calculated according to the intermediate tensor through a plurality of product accumulators. Therefore, frequent memory access of the intermediate tensor between the self-attention mechanism calculation array and the external memory is avoided, the calculation efficiency is improved, and the problems of intermediate tensor multiplexing and storage optimization are solved.
Owner:BYD SEMICON CO LTD

Systolic array having support for output sparsity

A processing apparatus is described herein that includes a general-purpose parallel processing engine comprising a matrix accelerator including one or more systolic arrays, at least one of the one or more systolic arrays comprising multiple pipeline stages, each pipeline stage of the multiple pipeline stages including multiple processing elements, the multiple processing elements configured to perform processing operations on input matrix elements based on output sparsity metadata. The output sparsity metadata indicates to the multiple processing elements to bypass multiplication for a first row of elements of a second matrix and multiply a second row of elements of the second matrix with a column of matrix elements of a first matrix.
Owner:INTEL CORP

Systolic array scheduling processing method, device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a systolic array scheduling processing method, device, equipment and medium, and the method comprises the steps: obtaining a data processing model, and compiling the data processing model into an initial processing task based on hardware architecture features of a systolic array, acquiring to-be-processed data and analyzing data characteristics of the to-be-processed data, monitoring a real-time operation state of the systolic array processing device, inputting the real-time operation state, the data characteristics and an initial processing task into a scheduling decision module to generate a scheduling strategy, executing the scheduling strategy to complete data processing, and collecting performance data; and updating optimization parameters of the compiling module and strategy generation parameters of the scheduling decision module based on the performance data. According to the method, the initial processing task is generated through compiling, the scheduling strategy is dynamically generated in combination with the data features and the running state, parameter updating is achieved through performance data feedback, self-adaptive closed-loop optimization of compiling and scheduling is formed, and the calculation efficiency and the energy efficiency ratio are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Machine learning training architecture for programmable devices

A programmable device may be configured to support machine learning training operations using matrix multiplication circuitry. In some embodiments, the multiplication is implemented on a systolic array. The systolic array includes an array of processing elements, each of which includes hybrid floating-point dot-product circuitry.
Owner:ALTERA CORP

An edge computing platform and system based on FPGA real-time target recognition detection

The application discloses an edge computing platform and system based on FPGA real-time target recognition detection, and relates to the field of microelectronic chips.The edge computing platform comprises an interconnected FPGA and DDR, and an ISP module, a pre-processing module, a VDMA module, an inference accelerator, a CPU and a character superposition module are arranged on the FPGA.The inference accelerator is used for deploying a preset convolutional neural network model, reading any frame of second image and weight parameter data about convolution kernels in the convolutional neural network model from the DDR, accelerating the execution of the algorithm of the convolutional neural network model by using a sphygmic array cluster module, and generating an inference result vector output to the DDR;the sphygmic array cluster module is integrated with a Winograd fast convolution algorithm and a multi-channel sphygmic array.Compared with the prior art, the application realizes the accelerated operation of the convolutional neural network model, and thus improves the inference efficiency.
Owner:GUANGDONG UNIV OF TECH

Systolic array matrix accelerator for graphics processing unit applications

One embodiment provides a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets coupled with the plurality of chiplet sockets, at least one of the plurality of chiplets including a plurality of processing elements and a matrix accelerator coupled with the plurality of processing elements, the matrix accelerator having circuitry to perform a matrix multiply accumulate operation on matrix data having a tiled memory layout.
Owner:INTEL CORP

Register file for systolic array

A processing apparatus includes a general-purpose parallel processing engine including a set of multiple processing elements including a single precision floating-point unit, a double precision floating point unit, and an integer unit; a matrix accelerator including one or more systolic arrays; a first register file coupled with a first read control circuit, wherein the first read control circuit couples with the set of multiple processing elements and the matrix accelerator to arbitrate read requests to the first register file from the set of multiple processing elements and the matrix accelerator; and a second register file coupled with a second read control circuit, wherein the second read control circuit couples with the matrix accelerator to arbitrate read requests to the second register file from the matrix accelerator and limit access to the second register file by the set of multiple processing elements.
Owner:INTEL CORP

Energy-efficient pre-encoded booth for stationary weights and activations

PendingUS20250377861A1Digital data processing detailsBooth encodingOperand
A neural network accelerator can perform energy-efficient multiply-and-accumulate operations of a neural network by Booth encoding a stationary operand, such as weights, before a compute phase. The Booth-encoding circuitry generates and stores Booth encoded multipliers in a Booth encoded multiplier storage and a precomputed compensation value representing a sum of the compensation bits of the Booth encoded multipliers in a Booth compensation storage. Per-cycle Booth encoding and compute of the sum of the compensation bits are avoided during multiply-accumulate operations because Booth encoding is applied to stationary operands. The Booth encoder can be located at the periphery where the multiplicands are loaded onto the accelerator shared across multiple compute columns and / or tiles to amortize the Booth encoder area overhead. The Booth encoder supports reconfigurable operand bit widths (e.g., 16-, 8-, 4-, and 2-bit). The approach is applicable to single-instruction-multiple data (SIMD) arrays, systolic arrays, and analog / digital compute-in-memory arrays.
Owner:INTEL CORP

Low latency matrix multiply unit

Methods, systems, and apparatus for a matrix multiply unit implemented as a systolic array of cells are disclosed. The matrix multiply unit may include cells arranged in columns of the systolic array. Two chains of weight shift registers per column of the systolic array are in the matrix multiply unit. Each weight shift register is connected to only one chain and each cell is connected to only one weight shift register. A weight matrix register per cell is configured to store a weight input received from a weight shift register. A multiply unit is coupled to the weight matrix register and configured to multiply the weight input of the weight matrix register with a vector data input in order to obtain a multiplication result.
Owner:GOOGLE LLC

Convolutional neural network hardware acceleration method and system

The invention relates to the technical field of hardware acceleration, and provides a convolutional neural network hardware acceleration method and system, and the method comprises the steps: obtaining a convolutional neural network instruction which comprises an img2col instruction, a systolic array multiplication instruction and a col2img instruction; analyzing an operation instruction code and a function code field of each instruction, determining a target register, a source register and calculation parameters, determining an instruction function and generating a control signal; data to be operated are taken out from the source register and transmitted to the corresponding operation module together with the control signal and the calculation parameter to be calculated, and then the data are written back to the target register; through hardware optimization strategies such as multi-level address superposition, systolic array data multiplexing, double-buffer window construction and a parallel comparison tree, extra power consumption caused by multiple times of data loading and storage and intermediate result calculation is reduced, and data handling overhead, calculation delay and control complexity are greatly reduced.
Owner:SHANDONG UNIV

Method and system for improving calculation efficiency of end-side convolution acceleration engine and storage medium

The invention discloses a method and system for improving the calculation efficiency of an end-side convolution acceleration engine and a storage medium, and belongs to the technical field of convolution acceleration engines, and the method comprises the steps: determining the weight of the end-side convolution acceleration engine, the size of an input feature map and the size of an output feature map under different conditions; correspondingly dividing an on-chip memory, which is divided into a plurality of memory blocks with preset sizes in advance, of the end-side convolution acceleration engine into three on-chip memory parts according to the sizes of the weight, the input feature graph and the output feature graph and the priorities of the weight, the input feature graph and the output feature graph from high to low, the weights, the input feature maps and the output feature maps are independently stored; and carrying out convolution calculation in a systolic array mode by using the independently stored weights, the input feature map and the output feature map. According to the specification of the current convolution calculation, the features and the weights are dynamically split, and the on-chip memory is dynamically allocated, so that the number of repeated data carrying times is reduced, and the calculation efficiency of the end-side convolution acceleration engine is improved.
Owner:CCORE TECH CO LTD

Low latency matrix multiply unit

Methods, systems, and apparatus for a matrix multiply unit implemented as a systolic array of cells are disclosed. Each cell of the matrix multiply includes: a weight matrix register configured to receive a weight input from either a transposed or a non-transposed weight shift register; a transposed weight shift register configured to receive a weight input from a horizontal direction to be stored in the weight matrix register; a non-transposed weight shift register configured to receive a weight input from a vertical direction to be stored in the weight matrix register; and a multiply unit that is coupled to the weight matrix register and configured to multiply the weight input of the weight matrix register with a vector data input in order to obtain a multiplication result.
Owner:GOOGLE LLC

SPAD imaging data compression system suitable for extremely low illumination

The invention discloses an SPAD imaging data compression system suitable for extremely low illumination, and the system comprises a data rearrangement module which is used for arranging one-dimensional discrete data corresponding to a plurality of pixels received by an SPAD array into multiple rows of data, and inputting the multiple rows of data corresponding to each pixel into a pulsation calculation module; the pulsation calculation module comprises a plurality of parallel pulsation arrays and an addition unit, and the addition unit and each pulsation array are used for performing wavelet decomposition calculation on the multi-row data of the corresponding pixel to obtain a wavelet decomposition result of the multi-row data of each pixel; carrying out summation on the wavelet decomposition results of all the pixels in the fixed neighborhood range of each pixel to obtain a summation result of each pixel; and the sliding window interception module is used for carrying out sliding window interception operation on the summation result of each pixel so as to obtain compressed data of each pixel. According to the invention, the processing time of the SPAD hardware system can be reduced, and the processing efficiency of the SPAD hardware system can be improved.
Owner:XIDIAN UNIV +1

A fast code generation device supporting the generation of fusion operators

A fast code generation device supporting the generation of fusion operators, belonging to the technical field of deep learning. The present invention includes: an LDM area division module for functionally partitioning the local storage space according to the network size parameters input by the upper-layer framework; a fusion operator address configuration module for defining the addresses of the input, output, and intermediate result data in the operator in the functional partition according to the type of fusion operator input by the upper-layer framework; a fusion operator data interaction module providing function interfaces for asynchronous memory access between the local and the main memory, and between the local and the local; a SIMD fusion operator calculation module for fusing the operator according to the addresses generated by the fusion operator address configuration module; and a systolic array instruction configuration module for configuring the instructions for driving the systolic array to perform calculations. The present invention can effectively reduce the code error rate, improve the code generation efficiency, and simplify the debugging process.
Owner:JIANGNAN INST OF COMPUTING TECH

A vector processor and processing method supporting multiple precision calculation and dynamic configuration

The vector processor and the data processing method provided by the application add a systolic array acceleration unit in a processor channel to realize the calculation between vectors. The storage unit on the original architecture is fully utilized, the data throughput is increased, the calculation between more vector data is realized, the acceleration effect of the systolic array accelerator is fully utilized, and the utilization rate of the calculation is greatly improved. The systolic array accelerator can support multi-precision and ultra-low bit quantization calculation, improve the efficiency of vector calculation, and the parallelism and scalability of the vector processor can greatly improve the data calculation density, thereby effectively improving the computing power.
Owner:NANJING UNIV

Clocking a systolic array on both edges of a clock signal

The present disclosure is directed to an apparatus that includes a systolic array having an initial systolic stage that is clocked at a first edge of a clock signal. The apparatus further includes control logic configured to clock a first subset of a plurality of additional systolic stages of the systolic array at the first edge of the clock signal. The control logic is further configured to clock a second subset of the plurality of systolic stages at a second edge of the clock signal.
Owner:QUALCOMM INC