Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

18 results about "Multiply–accumulate operation" patented technology

In computing, especially digital signal processing, the multiply–accumulate operation is a common step that computes the product of two numbers and adds that product to an accumulator. The hardware unit that performs the operation is known as a multiplier–accumulator (MAC, or MAC unit); the operation itself is also often called a MAC or a MAC operation. The MAC operation modifies an accumulator a: a←a+(b×c) When done with floating point numbers, it might be performed with two roundings (typical in many DSPs), or with a single rounding.

Data processing method, processor, chip, and electronic device

The present disclosure relates to a data processing method, a processor, a chip, and an electronic device. The method comprises: on the basis of an obtained control instruction, a control logic unit sequentially reads, from a memory to a dot product unit array, loop tiling data of data to be processed; the dot product unit array performs multiply-accumulate operation on the loop tiling data of the data to be processed that is received each time, and determines a loop tiling result of the data to be processed that is received each time; and on the basis of a plurality of loop tiling results obtained from the dot product unit array, the control logic unit determines a logic operation result of the data to be processed. The embodiments of the present disclosure can convert, into the reading and logic operation of multiple pieces of loop tiling data of the data to be processed, the reading and logic operation of the data to be processed, so that data with larger size can be processed under the condition that the hardware resources of the processor are not changed, and the pressure on the storage bandwidth is reduced.
Owner:MOORE THREADS TECH CO LTD

Memory device and computing method

PendingCN121096388ADigital data processing detailsDigital storageBit lineMultiply–accumulate operation
The invention provides an in-memory computing memory device. The memory device includes: a first weight group generating a first input weight product current on a first common bit line according to one of a plurality of inputs; the second weight group generates a second input weight product current on a second common bit line according to one of the inputs. The first common bit line and the second common bit line output the first input weight product current and the second input weight product current to a first differential analog-to-digital converter. The first differential analog-to-digital converter outputs a product accumulation operation result according to the first input weight product current and the second input weight product current.
Owner:MACRONIX INTERNATIONAL CO LTD

Hardware-software collaborative dnn operation acceleration method based on bit-level sparsity and fpga architecture

The application discloses a kind of based on bit-level sparsity software and hardware collaborative acceleration method and FPGA architecture, the method includes: based on bit-level serial arithmetic unit composition target array, based on coarse-grained coding method, operand is converted into the encoding representation for bit-level sparse calculation;Bit-level serial arithmetic unit is designed in cooperation with coarse-grained coding method;Based on lookup table resource characteristics, bit-level serial arithmetic unit is optimized;Based on target array, the product accumulation operation for encoding representation is executed;Based on distributed controller, through micro load balancing communication strategy, data flow and calculation scheduling in target array are managed.Based on coarse-grained coding method and the use of bit-level serial arithmetic unit, while maintaining the accuracy of calculation, information processing density and hardware resource utilization are improved;Based on micro load balancing communication strategy, the uncertainty of calculation time and load imbalance are improved, and then the product accumulation operation efficiency and resource utilization of FPGA platform are improved.
Owner:UNIV OF SCI & TECH OF CHINA

Error upper bound device, determination method, medium, terminal and program product suitable for dynamic precision floating point multiply accumulate operation

ActiveCN121523639BPathPingMultiply–accumulate operation
The application provides an error upper bound device, a determination method, a medium, a terminal and a program product suitable for dynamic precision floating point multiply-accumulate operation, comprising: a feature acquisition module for acquiring a truncation compensation factor of each operand and an accumulator tolerance factor; a local threshold allocation module for allocating a local error threshold for the current multiply-accumulate operation according to a preset global error tolerance, a remaining term number, a remaining budget, an energy consumption mode and the accumulator tolerance factor; a bit budget formula calculation module for calculating the actual precision supply and precision demand of each path respectively; a determination and feedback module for determining according to the actual precision supply and precision demand calculated by each path; if each path passes, a precision pass signal is fed back; otherwise, a precision upgrade mechanism is triggered. The application can give a deterministic error upper bound for the truncation error of each multiply-accumulate operation, and ensure that the global precision target is met.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

Generalized acceleration of matrix multiply accumulate operations

A method, computer readable medium, and processor are disclosed for performing matrix multiply and accumulate (MMA) operations. The processor includes a datapath configured to execute the MMA operation to generate a plurality of elements of a result matrix at an output of the datapath. Each element of the result matrix is generated by calculating at least one dot product of corresponding pairs of vectors associated with matrix operands specified in an instruction for the MMA operation. A dot product operation includes the steps of: generating a plurality of partial products by multiplying each element of a first vector with a corresponding element of a second vector; aligning the plurality of partial products based on the exponents associated with each element of the first vector and each element of the second vector; and accumulating the plurality of aligned partial products into a result queue utilizing at least one adder.
Owner:NVIDIA CORP

Neural network device performing floating point operation and operating method thereof

A neural network device performing floating point operations and an operating method thereof are provided. The neural network device performs a multiply-accumulate (MAC) operation for a product of a fraction of a weight and an input activation in a block floating point format by using an analog crossbar array, performs an addition operation for a shared exponent of the weight and the input activation in the block floating point format by using a digital computing circuit, and outputs a partial sum of a floating point output activation by combining a result of the MAC operation with a result of the addition operation.
Owner:SAMSUNG ELECTRONICS CO LTD

Error upper bound device suitable for dynamic precision floating point multiply-accumulate operation, judgment method, medium, terminal and program product

ActiveCN121523639ADigital data processing detailsEnergy efficient computingPathPingMultiply–accumulate operation
The invention provides an error upper bound device suitable for dynamic precision floating point multiply-accumulate operation, a judgment method, a medium, a terminal and a program product. The error upper bound device comprises a feature acquisition module, a calculation module and a calculation module, wherein the feature acquisition module is used for acquiring a truncation compensation factor of each operand and an accumulator tolerance factor; the local threshold value distribution module is used for distributing a local error threshold value for the current multiply-accumulate operation according to the preset global error tolerance, the residual term number, the residual budget, the energy consumption mode and the accumulator tolerance factor; the bit budget calculation module is used for calculating the plurality of error paths to obtain the actual precision supply and precision demand of each path; the judgment and feedback module is used for judging according to the actual precision supply and precision demand calculated by each path; if each path passes, feeding back a precision passing signal; otherwise, triggering a precision upgrading mechanism. According to the method, a deterministic error upper bound can be given for the truncation error of each multiply-accumulate operation, and the global precision target is ensured to be met.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

A method for implementing a large-scale optoelectronic reservoir computing system

ActiveCN116578163BAlgorithmMultiply–accumulate operation
The present application belongs to the technical field of reservoir computing, and particularly relates to a method for implementing a large-scale optoelectronic reservoir computing system. The method comprises: performing matrix sparsification connection to form a sparse connection matrix in the form of a lower triangular matrix composed of basic matrix calculation units A and B; then performing rank reduction operation on the sparse connection matrix, so that the product of the original sparse matrix and the input signal is converted into the product form of each basic matrix unit after splitting and the input signal after splitting, and the dimension of each subunit is reduced to 1 / 2 of the original dimension; the above splitting operation is repeated until the splitting reaches the scale supported by the optical chip unit; then the reservoir is trained and tested to obtain a weight matrix calculation prediction value. The multiplication operation shares the same chip, and the multiplication and accumulation operation of a matrix of any scale can be completed by multiplexing the basic scale optical computing chip. The present application achieves ideal effects in the application of communication signal post-equalization, signal recognition, etc.
Owner:FUDAN UNIVERSITY

Adaptive precision floating point multiply accumulate operation apparatus, method, medium, terminal and program product

The application provides a self-adaptive precision floating-point multiply-accumulate operation device, method, medium, terminal and program product, comprising: an unpacking and feature extraction module which unpacks and performs feature extraction on the received floating-point number to be operated; a precision level prediction module which predicts an initial precision level according to the input key features and the pre-set energy consumption mode and error threshold; a precision control module which generates an effective precision bit number control signal according to the initial precision level; a multiplication module which performs multiplication operation on the floating-point number to be operated according to the received effective precision bit number control signal; an error upper bound module which calculates the error upper bound; and compares the error upper bound with the error threshold; and a fusion accumulation module which accumulates the final multiplication result after normalization and rounding operation when the error upper bound is less than or equal to the error threshold, and outputs the operation result. The application can improve operation efficiency, reduce energy consumption and cost, and enhance numerical stability and reliability.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

Intelligent processing unit and convolution operation method

The invention discloses an intelligent processing unit and a convolution operation method, and belongs to the technical field of convolution operation, and the intelligent processing unit comprises a memory which is used for storing first input data and first weight data; the quantization circuit is used for quantizing the first input data to generate a plurality of second input data and an input data displacement, and is used for quantizing the first weight data to generate a plurality of second weight data and a weight data displacement; the product accumulation circuit is used for performing product accumulation operation on the second input data and the second weight data to generate an intermediate result; the displacement calculation circuit is used for generating an intermediate result displacement based on the input data displacement and the weight data displacement; the order accumulation circuit is used for performing order accumulation operation on the intermediate result and an intermediate accumulation result based on the intermediate result displacement so as to generate a final accumulation result; and the data type conversion circuit is used for converting the final accumulation result to generate output data.
Owner:SIGMASTAR TECH LTD

Computing resource management method and apparatus, electronic device, and storage medium

ActiveCN121349709BResource allocationDigital data processing detailsMultiply–accumulate operationResource management
The application relates to the technical field of computers and provides a kind of computing resource management method, device, electronic equipment and storage medium, wherein the method comprises: in the case of performing matrix product accumulation operation, the working state of the first memory is set to the resource occupation state;The terminal write instruction of performing matrix product accumulation operation is used to write the final operation result of matrix product accumulation operation into the second memory, and the working state of the first memory is changed from the resource occupation state to the normal state.The method couples the state change operation with the write instruction of the final operation result through the terminal write instruction, ensures that the occupied first memory resource will be reliably released when the matrix product accumulation operation task successfully outputs the final operation result.Thereby, the resource occupation risk caused by incomplete recovery is effectively avoided, and the reliability and effectiveness of the computing resource management are improved.
Owner:SHANGHAI BIREN TECH CO LTD

Software and hardware collaborative DNN operation acceleration method based on bit-level sparsity and FPGA architecture

The invention discloses a software and hardware collaborative acceleration method based on bit-level sparsity and an FPGA architecture, and the method comprises the steps: forming a target array based on bit-level serial arithmetic units, and converting operands into coded representation for bit-level sparse calculation based on a coarse-grained coding method; the bit-level serial arithmetic unit is obtained by collaborative design with a coarse-grained coding method; optimizing the bit-level serial arithmetic unit based on lookup table resource characteristics; performing a product accumulation operation for the encoded representation based on the target array; and based on the distributed controller, managing data flow and calculation scheduling in the target array through a micro load balancing communication strategy. A coarse-grained coding method and a bit-level serial arithmetic unit are used, so that the information processing density and the hardware resource utilization rate are improved while the calculation precision is kept; based on a microcosmic load balancing communication strategy, the calculation time uncertainty and the load imbalance are improved, and then the multiplication and accumulation operation efficiency and the resource utilization rate of an FPGA platform are improved.
Owner:UNIV OF SCI & TECH OF CHINA

Adaptive precision floating point multiply-accumulate operation device and method, medium, terminal and program product

The invention provides a self-adaptive precision floating point multiply-accumulate operation device and method, a medium, a terminal and a program product, and the method comprises the steps: an unpacking and feature extraction module unpacks a received to-be-operated floating point number, and carries out the feature extraction; the precision gear prediction module predicts an initial precision gear according to the input key features and a preset energy consumption mode and an error threshold value; the precision control module generates an effective precision digit control signal according to the initial precision gear; the multiplication module carries out multiplication operation on the floating-point number to be operated according to the received effective precision digit control signal; the error upper bound module calculates an error upper bound; comparing the error upper bound with an error threshold value; and when the error upper bound is smaller than or equal to an error threshold value, the fusion accumulation module accumulates the final multiplication results after the normalization and rounding operations, and outputs an operation result. The method can improve the operation efficiency, reduce the energy consumption and cost, and enhance the numerical stability and reliability.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

Tensor processing circuitry

There is provided tensor processing circuitry comprising a plurality of dot-product units, each of which is configured to perform a multiply accumulate operation. A format conversion unit is configured to convert the format of a first data element before processing by the plurality of dot product units. The format conversion unit is configured to convert the first data element from a first data format to one or more data elements in a second floating point data format, the first data format being one of a plurality of data formats supported by the tensor processing circuitry and the second data format being a predefined floating-point data format in which data elements are input to the dot-product units. If the first data format is a higher precision data format than the second floating-point data format, the format conversion unit generates two or more data elements in the second floating-point data format.
Owner:ARM LTD

Multiply-accumulate (MAC) apparatus for in-memory computation

PendingCN122266416ADigital data processing detailsDigital storageCapacitanceMultiply–accumulate operation
Embodiments of the present disclosure relate to a multiply-accumulate (MAC) device for in-memory computing. A capacitive charge-coupled mode analog in-memory computing (CIM) bitcell array is configured to generate an analog output voltage corresponding to a multiply-accumulate (MAC) operation result using multi-bit weights. The analog output voltage is input to a dual-mode activation module that is selectively capable of operating in a deep neural network (DNN) mode and a spiking neural network (SNN) mode. The activation module includes a sample-and-hold (S&H) circuit, a comparator, a digital-to-analog converter (DAC) that can be reconfigured depending on the selected mode of the DNN mode and the SNN mode.
Owner:NOKIA NETWORKS OY

Computing resource management method and device, electronic equipment and storage medium

ActiveCN121349709AResource allocationDigital data processing detailsMultiply–accumulate operationResource management
The invention relates to the technical field of computers, and provides a computing resource management method and device, electronic equipment and a storage medium, and the method comprises the steps: setting the working state of a first memory as a resource occupation state under the condition of executing a matrix product accumulation operation; the terminal write-in instruction is used for executing the matrix product accumulation operation and is used for writing a final operation result of the matrix product accumulation operation into the second memory and converting the working state of the first memory from the resource occupation state to the conventional state. According to the method, state transition operation is coupled with a writing instruction of a final operation result through a terminal writing instruction, and it is ensured that when a matrix product accumulation operation task successfully outputs the final operation result, occupied first storage resources are reliably released. Therefore, the resource occupation risk caused by incomplete recovery is effectively avoided, and the reliability and effectiveness of computing resource management are improved.
Owner:SHANGHAI BIREN TECH CO LTD

Tensor processing circuitry

Tensor processing circuitry 17 comprises a plurality of dot-product units 100, each of which is configured to perform a multiply accumulate operation. A format conversion unit (20, Fig. 2) is configur
Owner:ARM LTD