Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

27 results about "Matrix multiplier" patented technology

Matrix multiplier, chip, device, data processing method, medium and product

The invention discloses a matrix multiplier, a chip, electronic equipment, a data processing method, a computer readable storage medium and a computer program product, and belongs to the field of artificial intelligence. The matrix multiplier comprises a multiplier array, a data preprocessing module and an accumulator; the multiplier array is configured to perform multiplication operation on an input first matrix and a second matrix; the data preprocessing module is configured to preprocess the third matrix to obtain a mask matrix; the accumulator is configured to add the multiplication result of the first matrix and the second matrix and the mask matrix; wherein the first matrix is an M * K matrix, the second matrix is a K * N matrix, the third matrix is an M * N matrix, and M, N and K are all integers greater than or equal to 1. According to the method, the mask process is fused into the matrix multiplication operation, so that the matrix multiplication operation and the mask adding operation can be synchronously carried out, and the calculation rate of the Attention (attention) operation is effectively improved.
Owner:NANJING TIANSHU ZHIQI TECHNOLOGY CO LTD

Area efficient 3D NAND-based vector-matrix multiplier circuit with common-mode current cancellation

To reduce the area requirements for sensing circuits of 3D NAND-based vector-matrix multiplication circuitry where weight values for a neural network are stored differentially as current levels on pairs of memory cells, techniques are presented for reducing the common mode current levels during sensing operations. When discharging a first capacitor through a first of a memory cell of a pair of memory cells storing a weight value by a first bit line and discharging second capacitor through a second memory cell of the pair by a second bit line, a reference current is applied to the bit lines. The product of a weight value with an input vector values is then determined by comparing the voltage levels on the two capacitors. The use of the reference current reduces the amount of voltage swing in the two capacitors, reducing the size requirements for the capacitors.
Owner:SANDISK TECHNOLOGIES LLC

Field programmable gate array architecture optimized for machine learning applications

Systems and methods for a new field programmable gate array (FPGA) architecture that is optimized for machine learning (ML) applications are provided. Such ML applications can specifically include, for example, artificial neural networks and deep neural networks. Various embodiments enable the design of faster and more power efficient hardware accelerators for machine learning algorithms, compared to existing FPGAs in the market. This is made possible by hard systolic matrix multiplier blocks, hard activation blocks and soft ML-centric configurable logic blocks. The matrix multiplier blocks are connected to field programmable interconnect resources to enable creation of larger matrix multipliers. The hard matrix multipliers and the hard activation blocks have programmable interconnects between them and neighboring memory or compute blocks on the device.
Owner:BOARD OF RGT THE UNIV OF TEXAS SYST

Methods and apparatus for vector lane matrix multiplication

PendingUS20260111391A1Digital computer detailsProgram controlBinary multiplierMatrix multiplier
Systems, apparatus, articles of manufacture, and methods are disclosed. An example apparatus includes a Vector Processor Unit (VPU) comprising: first vector lane circuitry including first matrix multiplier circuitry; second vector lane circuitry including second matrix multiplier circuitry; and interconnect circuitry to connect the first vector lane circuitry and the second vector lane circuitry in a ring structure.
Owner:OPENCHIP & SOFTWARE TECHNOLOGIES SL

Structured sparse matrix multiplier realized based on FPGA primitive

The invention discloses a structured sparse matrix multiplier realized based on FPGA primitives, and belongs to the technical field of FPGA hardware acceleration and deep learning computing architecture. The multiplier is composed of a register area and a multiplication and addition area, and the core design thought is that a circuit is built in a customized mode by directly calling bottom layer physical resources based on FPGA primitives; the register area stores a dense matrix B and supports parallel reading of elements by utilizing the characteristic that LUT in an SLICEM can be configured to be double SRL16E; the multiplication and addition area refers to a partial product generation unit based on the 4-Booth algorithm and an improved GPC (4: 2) compressor structure, and efficient generation and rapid accumulation of partial products are achieved. According to the method, redundant loss caused by high-level HDL logic synthesis is avoided through precise physical resource binding of primitives, fine wiring constraint and structured sparse data characteristic adaptation.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Throughput optimized 3D NAND-based vector-by-matrix multiplier circuit

To improve the speed of the 3D NAND-based vector-matrix multiplication circuitry, the circuit is designed such that the charge is accumulated on a slave capacitor which is not directly connected to the array. The master capacitor is reset in each cycle, reducing the maximum swing on the bit lines and, hence, reduce the capacitance size. The reduction of the capacitance size will allow the vector-matrix multiplication to run much faster because of smaller interconnect parasitics. A second set of aspects is based on modification of timings of each operation phase. Rather than equal time slots dedicated to different operation phases (e.g., the integration and scaling), the circuit is modified such the masking and scaling phases would be executed much faster as they typically have a much faster time constant.
Owner:SANDISK TECHNOLOGIES LLC

Extensible 2*2 processing unit matrix multiplier and Transform acceleration method

The invention discloses an extensible 2 * 2 processing unit matrix multiplier and a Transform acceleration method, when a control module executes an OS data stream mode, each processing unit performs multiplication and addition processing according to a received activation value and a weight value to obtain a first multiplication and addition result; each processing unit performs multiplication of an activation value and a weight value according to the processing unit interconnected with the processing unit to obtain a first operation result, and the first operation result is accumulated with the first multiply-add result to obtain a product result of two matrixes; when the control module executes the WS data stream mode, each processing unit performs multiplication and addition processing according to the received activation value and weight value to obtain a second multiplication and addition result; and each processing unit locks the weight value to be unchanged, performs multiplication operation with the activation value according to the processing unit interconnected with the processing unit to obtain a second operation result, and accumulates the second operation result with the second multiply-add result to obtain a matrix multiplication result based on the fixed weight, so that the efficiency ratio and the flexibility are improved.
Owner:XIDIAN UNIV

Implementation method, device and medium of a general fixed-point matrix multiplier based on FPGA high-performance computing architecture

This invention discloses a method, apparatus, and medium for implementing a general-purpose fixed-point matrix multiplier based on a high-performance computing architecture in an FPGA. The method includes: designing matrix partitioning strategies at different levels based on the parallelism of the AI ​​engine array resources on the Versal ACAP platform; data scheduling and multiplexing based on the data packet stream and data packet exchange of the AI ​​engine and AXI stream, accelerating the kernel through matrix partitioned multiplication, and designing data scheduling and multiplexing strategies; and implementing a high-throughput vectorized matrix multiplication pipeline on the AI ​​engine vector processor. This invention achieves multi-level partitioning of matrix multiplication based on the AXI stream transport protocol and AI engine array on the Versal ACAP platform, enabling efficient utilization of hardware resources, effectively improving data reuse rate, and achieving high parallelism, while achieving high computational speed under the high-speed clock of the AI ​​engine. This invention can be widely applied in the field of high-performance computing.
Owner:SOUTH CHINA UNIV OF TECH

Scalable 2x2 processing unit matrix multiplier and transform acceleration method

The application discloses an extensible 2*2 processing unit matrix multiplier and a Transformer acceleration method, when a control module executes an OS data flow mode, each processing unit respectively performs multiplication and addition processing according to received activation values and weight values, and obtains a first multiplication and addition result; each processing unit respectively performs multiplication operation of activation values and weight values according to processing units interconnected with the processing unit, and obtains a first operation result, and the first operation result is accumulated with the first multiplication and addition result to obtain a two-matrix product result; when the control module executes a WS data flow mode, each processing unit respectively performs multiplication and addition processing according to received activation values and weight values, and obtains a second multiplication and addition result; each processing unit locks the weight values unchanged, and respectively performs multiplication operation of the activation values according to the processing units interconnected with the processing unit, and obtains a second operation result, and the second operation result is accumulated with the second multiplication and addition result to obtain a matrix multiplication result based on fixed weight, and the efficiency ratio and flexibility are improved.
Owner:XIDIAN UNIV

Method and systems for permuted diagonal computing with ultrashort pulses

Systems and methods for optical matrix multiplication are disclosed. An optical matrix multiplier includes an optical source configured to produce a sequence of optical pulses based on an input vector. The multiplier further includes a fanout module configured to receive a sequence of optical pulses and produce multiple optical pulse sequences. The multiplier further includes multiple delay lines, each delay line configured to apply a delay to an associated one of the optical pulse sequences. The multiplier further includes multiple modulators, each modulator configured to modulate one of the delayed optical pulse sequences. The multiplier further includes at least one accumulator configured to sum at least one modulated optical pulse sequence.
Owner:UNIVERSITY OF ROCHESTER

Optical matrix multiplier

The invention discloses an optical matrix multiplier, which comprises a laser light source, a grating coupler, a BTO electro-optical modulation array, an MZI waveguide array, a photoelectric detector array and a control unit, and is characterized in that the laser light source generates coherent light with stable wavelength, and the coherent light enters a BTO waveguide after being coupled by the grating coupler and then enters the MZI waveguide array through multi-stage beam splitting; the excellent electro-optical effect of a BTO material is utilized, phase modulation is achieved under electrode driving so as to load matrix element information, and then linear operation of matrix multiplication is completed through interference superposition of an MZI waveguide array. And the output optical signal is converted into an electric signal by the photoelectric detector array and is processed by the control unit to obtain an operation result. The system is compact in structure, high in modulation linearity, high in response speed, low in power consumption and capable of achieving multi-channel parallel optical calculation. Compared with a traditional silicon-based or lithium niobate modulator, higher integration density and stability are achieved, and the calculation efficiency of matrix multiplication in the fields of artificial intelligence, scientific calculation, signal processing and the like is remarkably improved.
Owner:NANKAI UNIV

General purpose parallel matrix multiplier based on reconfigurable computation

The application discloses a general parallel matrix multiplier based on reconstruction calculation. The multiplier comprises a matrix reconstruction module, a multiplication module, a compression module, a shift module and an accumulation module. After matrix data is reconstructed into multiple groups of single-bit data according to bit positions by the matrix reconstruction module, the single-bit data is multiplied in the multiplication module; the multiplication result is compressed by the compression module. The compressor in the compression module can compress six data simultaneously. Finally, the compressed result is shifted by the shift module and accumulated by the accumulation module to obtain the multiplication result of the matrix data. According to the matrix multiplier provided by the application, the compression efficiency can be improved by compressing data by the compression module composed of the compressor; meanwhile, the compressor can also perform partial shift calculation and undertake part of the responsibility of the subsequent shift module, thereby reducing the area cost of the matrix multiplier; by decomposing the matrix data according to bit positions, the calculation of signed numbers and unsigned numbers can be realized.
Owner:XIDIAN UNIV

Customized temporary memory for partial dot product reduction

Aspects of the disclosed technology include techniques and mechanisms for partial dot product reduction using customized temporary memory. The custom temporary memory may be a dedicated memory that is dedicated to receive and store the partial dot product determined by the matrix multiplier unit. Each partial dot product may correspond to a tile of a resulting matrix, where the resulting matrix is a product of matrix multiplications that may use a first matrix representing a user query as a left operand, and a second matrix representing a user query as a right operand. And using a second matrix representing a trained model containing data available to respond to the user query as a right side operand. The custom register memory may append tiles determined by matrix multiplication, where the appended tiles may create a resulting matrix. Customized temporary memory may write the resulting matrix to general purpose memory, where the resulting matrix may be used to respond to a user query.
Owner:GOOGLE LLC

Matrix multiplier execution method and device, equipment and storage medium

The invention provides an execution method and device of a matrix multiplier, equipment and a storage medium, relates to the technical field of artificial intelligence, and is suitable for executing the matrix multiplier in parallel through a plurality of calculation thread bundle groups, and the method comprises the following steps: through the same carrying thread bundle group, matrix data required by the plurality of calculation thread bundle groups to execute the matrix multiplier is converted into matrix data required by the plurality of calculation thread bundle groups to execute the matrix multiplier; carrying to an on-chip cache from a video memory according to a loading sequence; wherein the loading sequence is that the preorder data required by the plurality of calculation thread beam groups are loaded in series, and the subsequent data required by each calculation thread beam group are continuously loaded; and respectively scheduling the calculation units through the plurality of calculation thread bundle groups, and carrying out matrix multiplication calculation on the respective required matrix data acquired from the on-chip cache. By continuously loading the subsequent data required by the same calculation thread bundle group, one calculation thread bundle group starts to calculate the subsequent steps by using the calculation unit more quickly, and the result of the calculation thread bundle group is more obviously staggered from the result of the other calculation thread bundle group, so that the calculation unit is prevented from being scrambled.
Owner:SHANGHAI BIREN TECH CO LTD

Method and apparatus for a vector computing device

A method for a vector computing device, for example a vector-matrix multiplier, comprising: supplying a first set of a first input quantity, for example in the form of a bit vector, to the vector computing device; charging a capacitor device with a first output current, which characterizes a product, for example a scalar product, of a multiplication of the first set of the first input quantity with a second input quantity by means of the vector computing device; partially discharging the capacitor device by a predefinable amount; repeating at least one of the aspects of supplying and charging, optionally also of discharging, using at least a second set of the first input quantity.
Owner:ROBERT BOSCH GMBH

Storage and calculation integrated peripheral circuit device based on 3D VRRAM and matrix calculation method

PendingCN122086353AConvenient for Embedded ApplicationsImprove storage densityDigital data processing detailsDigital storageMatrix additionBinary multiplier
The invention relates to a storage and calculation integrated peripheral circuit device based on a 3D VRRAM and a matrix calculation method, belongs to the technical field of memories, and solves the problem that the structure of an existing two-dimensional resistive random access memory array is not suitable for a neural network with a high calculation power demand. The device comprises a matrix multiplier for multiplying an input data matrix by a weight matrix; the input ends of the analog-to-digital converters are connected to the output ends of the column control switches so as to convert the analog product result into a digital product result; the plurality of samplers are connected with the output end of the analog-to-digital converter so as to sample the digital product result; the output ends of the plurality of samplers are connected with the input end of the matrix adder through the matrix adder so as to realize digital product result shifting, then shifting data are accumulated, and an accumulation result is stored in a temporary register; and the updating register adds the accumulation result and the data in the updating register to obtain sum data, and updates the data in the updating register by using the sum data. And high-storage and high-computing-power-density operation is realized so as to be suitable for a neural network.
Owner:INST OF MICROELECTRONICS CHINESE ACAD OF SCI LTD

Systolic array matrix multiplier and operation method of systolic array matrix multiplier

To execute various sizes of matrix multiplication while suppressing increase in circuit scale.SOLUTION: A systolic array matrix multiplier includes a plurality of processing elements arranged in a matrix and executes matrix multiplication. Each of the plurality of processing elements includes a first holding unit which sequentially holds respective elements of a first matrix received from a first input terminal provided in one end side in a first direction, a first path which outputs an output of the first holding unit to a first output terminal provided in the other end side in the first direction, a second holding unit which sequentially holds respective elements of the first matrix received from a second input terminal provided in the other end side in the first direction, a second path which outputs an output of the second holding unit to a second output terminal provided in one end side in the first direction, a product-sum operator connected to the first path, a first selection unit which connects the first path or the first output terminal to the second path, and a second selection unit which connects the second path or an output of the first holding unit to the first path.SELECTED DRAWING: Figure 3
Owner:FUJITSU LTD

A file encryption apparatus

This invention provides a file encryption device, comprising: a file input pool, a file formatter, a matrix multiplier, an inverse matrix unit, a file output pool, a user password manager, an encryption matrix generator, a universal DES encryptor, an encryption matrix array, and a universal DES decryptor. This invention can meet the needs of people's daily electronic data for low security and large data volumes, avoiding the use of expensive professional data encryption hardware or software, and preventing low-cost data cracking. It provides a low-computation, low-latency, and low-cost encryption device based on a matrix transformation algorithm.
Owner:ZHONGKE FANYU (WUHAN) TECH CO LTD

Photonic blockchain based on optical proof-of-work

An apparatus for combined digital and optical processing of a cryptocurrency data block includes a digital processor that computes a hash vector from the cryptocurrency data block; a laser and splitter that produces optical input signals; optical modulators that binary phase-shift key modulate the optical input signals based on the hash vector; a photonic matrix multiplier circuit that performs an optically perform a discrete matrix-vector product operation on the modulated optical input signals to produce optical output signals, where the discrete matrix-vector product operation is defined by matrix elements limited to K discrete values, where 2≤K≤17; and photodetectors and comparators that perform optoelectronic conversions of the optical output signals to produce corresponding digital electronic output signals. The digital processor performs a second hash computation on an XOR result between the digital electronic output signals and the hash vector to produce a proof of work result.
Owner:THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV +1

Systolic array matrix multiplier and method of operating systolic array matrix multiplier

A systolic array matrix multiplier executes matrix multiplication and includes processing element which each includes: a first holder that retains each element of a first matrix received from a first input terminal provided on one end side; a first path that outputs an output of the first holder to a first output terminal provided on another end side; a second holder that retains each element of the first matrix received from a second input terminal provided on the another end side; a second path that outputs an output of the second holder to a second output terminal provided on the one end side; a product-sum operator coupled to the first path; a first selector that couples the first path or the second input terminal to the second path; and a second selector that couples the second path or the output of the first holder to the first path.
Owner:FUJITSU LTD

Matrix multiplier for transformer-based model training

This invention provides a matrix multiplier for training Transformer-type models, comprising an M-row, N-column systolic array. The systolic array is two-dimensional and consists of R-row, C-column interconnected processing units (PEs). Each PE includes one multiplier, one adder, two internal registers, one left-side multiplexer, and two right-side multiplexers. The left-side multiplexer can select whether the input to the multiplier comes from outside the PE or retains the input from the previous cycle. When retaining the input from the previous cycle, the PE maintains the WS data stream with weights. This invention designs a reconfigurable processing unit (PE) that can flexibly support multiple data streams at different stages and cycles of training and select the data source according to requirements.
Owner:NANJING UNIV

Thin film lithium niobate microring filter arrays, optical vector-matrix multipliers

The application provides a thin-film lithium niobate micro-ring filter array and an optical vector-matrix multiplier, and comprises a thin-film lithium niobate micro-ring filter array; thin-film lithium niobate micro-ring filters with a first target row number and a first target column number, thin-film lithium niobate input waveguides with the first target row number, and thin-film lithium niobate output waveguides with the first target row number, and each thin-film lithium niobate micro-ring filter is provided with a driving module; the driving module is used for adjusting the initial resonant wavelength of the corresponding thin-film lithium niobate micro-ring filter to a target resonant wavelength. The application utilizes the electro-optic effect characteristics of thin-film lithium niobate to change the resonant wavelength of the micro-ring filter. The optical vector-matrix multiplier integrated with the thin-film lithium niobate micro-ring filter array of the application has a small power consumption, less than 20 fJ per byte.
Owner:张江国家实验室 +1

Methods and systems for optical matrix calculation

ActiveUS12499173B2Quantum computersDigital dataPhotodetectorMatrix multiplier
Aspects relate to methods and systems for optical matrix calculation. An exemplary system includes at least a first light source configured to output at least a first optical output having a first wavelength, at least a second light source configured to output at least a second optical output having a second wavelength substantially different from the first wavelength, at least an optical modulator configured to modulate the at least a first optical output, at least an optical matrix multiplier configured to perform at least two matrix multiplications, a first matrix multiplication as a function of the first optical output and a second matrix multiplication as a function of the second optical output, and at least a photodetector configured to measure the at least a first optical output and the at least a second optical output.
Owner:SIPHOX INC

Vector computing device

A vector processor is disposed on a chip and contains a multi-port shared memory, a unit for performing horizontal operations, and interconnected scalar devices, each of which is configured to be capable of receiving a vector element from a matrix multiplication device. A scalar device contains scalar modules, two demultiplexers, and a unit for performing complex arithmetic operations. The technical result is an increase in the operating speed of a vector processor and a decrease in the chip area thereof.
Owner:AKTSIONERNOE OBSHCHESTVO SOFIT

Neural inference processing unit with flexible precision

ActiveCN114787823BDigital computer detailsInference methodsAlgorithmMatrix multiplier
Neural inference chips are provided. A neural core of a neural inference chip includes a vector-matrix multiplier; a vector processor; and an activation unit, which is operatively coupled to the vector processor. The vector-matrix multiplier, vector processor, and / or activation unit are adapted to operate at variable precision.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Time-space-wavelength-multiplexed photonic matrix multiplier

PCT designated stageWO2025198610A1Digital dataCoupling light guidesGratingMatrix multiplier
A photonic matrix multiplier multiplies an MxP first matrix with a PxN second matrix. The multiplier includes a splitter that receives light including M different wavelengths into M separate light beams. M first optical modulators modulate the M light beams with time-multiplexed row vectors of the first matrix. An arrayed waveguide grating router (AWGR) processes the output signals of the M first optical modulators and outputs N processed light signals. N second optical modulators modulate the N processed light signals with time-multiplexed column vectors of the second matrix. N demultiplexers separate the N second optical modulator output signals in light beams at each of the M different wavelengths. Integrators integrate demultiplexer output signals to obtain multiplied and accumulated signals.
Owner:CELESTIAL AI INC

Matrix multiplier cache

Techniques related to integrated circuits supporting matrix operations are disclosed. In various embodiments, an integrated circuit includes a dot product accumulation circuit including: a dot product circuit configured to determine a dot product of a first vector and a second vector; and an adder circuit coupled to an output of the dot product circuit and configured to add a result of the dot product to the accumulated value. The integrated circuit also includes an accumulator cache coupled to an input of the adder circuit and an output of the adder circuit. The accumulator cache is configured to provide the accumulated value to the adder circuit, and store the result of the addition as a subsequent accumulated value for a subsequent point accumulation addition operation.
Owner:APPLE INC