Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

17 results about "Matrix multiplier" patented technology

Area efficient 3D NAND-based vector-matrix multiplier circuit with common-mode current cancellation

To reduce the area requirements for sensing circuits of 3D NAND-based vector-matrix multiplication circuitry where weight values for a neural network are stored differentially as current levels on pairs of memory cells, techniques are presented for reducing the common mode current levels during sensing operations. When discharging a first capacitor through a first of a memory cell of a pair of memory cells storing a weight value by a first bit line and discharging second capacitor through a second memory cell of the pair by a second bit line, a reference current is applied to the bit lines. The product of a weight value with an input vector values is then determined by comparing the voltage levels on the two capacitors. The use of the reference current reduces the amount of voltage swing in the two capacitors, reducing the size requirements for the capacitors.
Owner:SANDISK TECHNOLOGIES LLC

Field programmable gate array architecture optimized for machine learning applications

Systems and methods for a new field programmable gate array (FPGA) architecture that is optimized for machine learning (ML) applications are provided. Such ML applications can specifically include, for example, artificial neural networks and deep neural networks. Various embodiments enable the design of faster and more power efficient hardware accelerators for machine learning algorithms, compared to existing FPGAs in the market. This is made possible by hard systolic matrix multiplier blocks, hard activation blocks and soft ML-centric configurable logic blocks. The matrix multiplier blocks are connected to field programmable interconnect resources to enable creation of larger matrix multipliers. The hard matrix multipliers and the hard activation blocks have programmable interconnects between them and neighboring memory or compute blocks on the device.
Owner:BOARD OF RGT THE UNIV OF TEXAS SYST

Methods and apparatus for vector lane matrix multiplication

PendingUS20260111391A1Digital computer detailsProgram controlBinary multiplierMatrix multiplier
Systems, apparatus, articles of manufacture, and methods are disclosed. An example apparatus includes a Vector Processor Unit (VPU) comprising: first vector lane circuitry including first matrix multiplier circuitry; second vector lane circuitry including second matrix multiplier circuitry; and interconnect circuitry to connect the first vector lane circuitry and the second vector lane circuitry in a ring structure.
Owner:OPENCHIP & SOFTWARE TECHNOLOGIES SL

Structured sparse matrix multiplier realized based on FPGA primitive

The invention discloses a structured sparse matrix multiplier realized based on FPGA primitives, and belongs to the technical field of FPGA hardware acceleration and deep learning computing architecture. The multiplier is composed of a register area and a multiplication and addition area, and the core design thought is that a circuit is built in a customized mode by directly calling bottom layer physical resources based on FPGA primitives; the register area stores a dense matrix B and supports parallel reading of elements by utilizing the characteristic that LUT in an SLICEM can be configured to be double SRL16E; the multiplication and addition area refers to a partial product generation unit based on the 4-Booth algorithm and an improved GPC (4: 2) compressor structure, and efficient generation and rapid accumulation of partial products are achieved. According to the method, redundant loss caused by high-level HDL logic synthesis is avoided through precise physical resource binding of primitives, fine wiring constraint and structured sparse data characteristic adaptation.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Throughput optimized 3D NAND-based vector-by-matrix multiplier circuit

To improve the speed of the 3D NAND-based vector-matrix multiplication circuitry, the circuit is designed such that the charge is accumulated on a slave capacitor which is not directly connected to the array. The master capacitor is reset in each cycle, reducing the maximum swing on the bit lines and, hence, reduce the capacitance size. The reduction of the capacitance size will allow the vector-matrix multiplication to run much faster because of smaller interconnect parasitics. A second set of aspects is based on modification of timings of each operation phase. Rather than equal time slots dedicated to different operation phases (e.g., the integration and scaling), the circuit is modified such the masking and scaling phases would be executed much faster as they typically have a much faster time constant.
Owner:SANDISK TECHNOLOGIES LLC

Scalable 2x2 processing unit matrix multiplier and transform acceleration method

The application discloses an extensible 2*2 processing unit matrix multiplier and a Transformer acceleration method, when a control module executes an OS data flow mode, each processing unit respectively performs multiplication and addition processing according to received activation values and weight values, and obtains a first multiplication and addition result; each processing unit respectively performs multiplication operation of activation values and weight values according to processing units interconnected with the processing unit, and obtains a first operation result, and the first operation result is accumulated with the first multiplication and addition result to obtain a two-matrix product result; when the control module executes a WS data flow mode, each processing unit respectively performs multiplication and addition processing according to received activation values and weight values, and obtains a second multiplication and addition result; each processing unit locks the weight values unchanged, and respectively performs multiplication operation of the activation values according to the processing units interconnected with the processing unit, and obtains a second operation result, and the second operation result is accumulated with the second multiplication and addition result to obtain a matrix multiplication result based on fixed weight, and the efficiency ratio and flexibility are improved.
Owner:XIDIAN UNIV

Method and systems for permuted diagonal computing with ultrashort pulses

Systems and methods for optical matrix multiplication are disclosed. An optical matrix multiplier includes an optical source configured to produce a sequence of optical pulses based on an input vector. The multiplier further includes a fanout module configured to receive a sequence of optical pulses and produce multiple optical pulse sequences. The multiplier further includes multiple delay lines, each delay line configured to apply a delay to an associated one of the optical pulse sequences. The multiplier further includes multiple modulators, each modulator configured to modulate one of the delayed optical pulse sequences. The multiplier further includes at least one accumulator configured to sum at least one modulated optical pulse sequence.
Owner:UNIVERSITY OF ROCHESTER

Optical matrix multiplier

The invention discloses an optical matrix multiplier, which comprises a laser light source, a grating coupler, a BTO electro-optical modulation array, an MZI waveguide array, a photoelectric detector array and a control unit, and is characterized in that the laser light source generates coherent light with stable wavelength, and the coherent light enters a BTO waveguide after being coupled by the grating coupler and then enters the MZI waveguide array through multi-stage beam splitting; the excellent electro-optical effect of a BTO material is utilized, phase modulation is achieved under electrode driving so as to load matrix element information, and then linear operation of matrix multiplication is completed through interference superposition of an MZI waveguide array. And the output optical signal is converted into an electric signal by the photoelectric detector array and is processed by the control unit to obtain an operation result. The system is compact in structure, high in modulation linearity, high in response speed, low in power consumption and capable of achieving multi-channel parallel optical calculation. Compared with a traditional silicon-based or lithium niobate modulator, higher integration density and stability are achieved, and the calculation efficiency of matrix multiplication in the fields of artificial intelligence, scientific calculation, signal processing and the like is remarkably improved.
Owner:NANKAI UNIV

General purpose parallel matrix multiplier based on reconfigurable computation

The application discloses a general parallel matrix multiplier based on reconstruction calculation. The multiplier comprises a matrix reconstruction module, a multiplication module, a compression module, a shift module and an accumulation module. After matrix data is reconstructed into multiple groups of single-bit data according to bit positions by the matrix reconstruction module, the single-bit data is multiplied in the multiplication module; the multiplication result is compressed by the compression module. The compressor in the compression module can compress six data simultaneously. Finally, the compressed result is shifted by the shift module and accumulated by the accumulation module to obtain the multiplication result of the matrix data. According to the matrix multiplier provided by the application, the compression efficiency can be improved by compressing data by the compression module composed of the compressor; meanwhile, the compressor can also perform partial shift calculation and undertake part of the responsibility of the subsequent shift module, thereby reducing the area cost of the matrix multiplier; by decomposing the matrix data according to bit positions, the calculation of signed numbers and unsigned numbers can be realized.
Owner:XIDIAN UNIV

Customized temporary memory for partial dot product reduction

Aspects of the disclosed technology include techniques and mechanisms for partial dot product reduction using customized temporary memory. The custom temporary memory may be a dedicated memory that is dedicated to receive and store the partial dot product determined by the matrix multiplier unit. Each partial dot product may correspond to a tile of a resulting matrix, where the resulting matrix is a product of matrix multiplications that may use a first matrix representing a user query as a left operand, and a second matrix representing a user query as a right operand. And using a second matrix representing a trained model containing data available to respond to the user query as a right side operand. The custom register memory may append tiles determined by matrix multiplication, where the appended tiles may create a resulting matrix. Customized temporary memory may write the resulting matrix to general purpose memory, where the resulting matrix may be used to respond to a user query.
Owner:GOOGLE LLC

Method and apparatus for a vector computing device

A method for a vector computing device, for example a vector-matrix multiplier, comprising: supplying a first set of a first input quantity, for example in the form of a bit vector, to the vector computing device; charging a capacitor device with a first output current, which characterizes a product, for example a scalar product, of a multiplication of the first set of the first input quantity with a second input quantity by means of the vector computing device; partially discharging the capacitor device by a predefinable amount; repeating at least one of the aspects of supplying and charging, optionally also of discharging, using at least a second set of the first input quantity.
Owner:ROBERT BOSCH GMBH

Storage and calculation integrated peripheral circuit device based on 3D VRRAM and matrix calculation method

PendingCN122086353AConvenient for Embedded ApplicationsImprove storage densityDigital data processing detailsDigital storageMatrix additionBinary multiplier
The invention relates to a storage and calculation integrated peripheral circuit device based on a 3D VRRAM and a matrix calculation method, belongs to the technical field of memories, and solves the problem that the structure of an existing two-dimensional resistive random access memory array is not suitable for a neural network with a high calculation power demand. The device comprises a matrix multiplier for multiplying an input data matrix by a weight matrix; the input ends of the analog-to-digital converters are connected to the output ends of the column control switches so as to convert the analog product result into a digital product result; the plurality of samplers are connected with the output end of the analog-to-digital converter so as to sample the digital product result; the output ends of the plurality of samplers are connected with the input end of the matrix adder through the matrix adder so as to realize digital product result shifting, then shifting data are accumulated, and an accumulation result is stored in a temporary register; and the updating register adds the accumulation result and the data in the updating register to obtain sum data, and updates the data in the updating register by using the sum data. And high-storage and high-computing-power-density operation is realized so as to be suitable for a neural network.
Owner:INST OF MICROELECTRONICS CHINESE ACAD OF SCI LTD

A file encryption apparatus

This invention provides a file encryption device, comprising: a file input pool, a file formatter, a matrix multiplier, an inverse matrix unit, a file output pool, a user password manager, an encryption matrix generator, a universal DES encryptor, an encryption matrix array, and a universal DES decryptor. This invention can meet the needs of people's daily electronic data for low security and large data volumes, avoiding the use of expensive professional data encryption hardware or software, and preventing low-cost data cracking. It provides a low-computation, low-latency, and low-cost encryption device based on a matrix transformation algorithm.
Owner:ZHONGKE FANYU (WUHAN) TECH CO LTD

Matrix multiplier for transformer-based model training

This invention provides a matrix multiplier for training Transformer-type models, comprising an M-row, N-column systolic array. The systolic array is two-dimensional and consists of R-row, C-column interconnected processing units (PEs). Each PE includes one multiplier, one adder, two internal registers, one left-side multiplexer, and two right-side multiplexers. The left-side multiplexer can select whether the input to the multiplier comes from outside the PE or retains the input from the previous cycle. When retaining the input from the previous cycle, the PE maintains the WS data stream with weights. This invention designs a reconfigurable processing unit (PE) that can flexibly support multiple data streams at different stages and cycles of training and select the data source according to requirements.
Owner:NANJING UNIV

Thin film lithium niobate microring filter arrays, optical vector-matrix multipliers

The application provides a thin-film lithium niobate micro-ring filter array and an optical vector-matrix multiplier, and comprises a thin-film lithium niobate micro-ring filter array; thin-film lithium niobate micro-ring filters with a first target row number and a first target column number, thin-film lithium niobate input waveguides with the first target row number, and thin-film lithium niobate output waveguides with the first target row number, and each thin-film lithium niobate micro-ring filter is provided with a driving module; the driving module is used for adjusting the initial resonant wavelength of the corresponding thin-film lithium niobate micro-ring filter to a target resonant wavelength. The application utilizes the electro-optic effect characteristics of thin-film lithium niobate to change the resonant wavelength of the micro-ring filter. The optical vector-matrix multiplier integrated with the thin-film lithium niobate micro-ring filter array of the application has a small power consumption, less than 20 fJ per byte.
Owner:张江国家实验室 +1

Vector computing device

A vector processor is disposed on a chip and contains a multi-port shared memory, a unit for performing horizontal operations, and interconnected scalar devices, each of which is configured to be capable of receiving a vector element from a matrix multiplication device. A scalar device contains scalar modules, two demultiplexers, and a unit for performing complex arithmetic operations. The technical result is an increase in the operating speed of a vector processor and a decrease in the chip area thereof.
Owner:AKTSIONERNOE OBSHCHESTVO SOFIT

Matrix multiplier cache

Techniques related to integrated circuits supporting matrix operations are disclosed. In various embodiments, an integrated circuit includes a dot product accumulation circuit including: a dot product circuit configured to determine a dot product of a first vector and a second vector; and an adder circuit coupled to an output of the dot product circuit and configured to add a result of the dot product to the accumulated value. The integrated circuit also includes an accumulator cache coupled to an input of the adder circuit and an output of the adder circuit. The accumulator cache is configured to provide the accumulated value to the adder circuit, and store the result of the addition as a subsequent accumulated value for a subsequent point accumulation addition operation.
Owner:APPLE INC