Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

28 results about "Adder tree" patented technology

Alignment in hardware accelerators

Systems, apparatuses, and methods are disclosed for improved matrix–vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values in floating point formats. The CIM macro has a functional block configured to align mantissa bits of primitive products between the activation values and the weight values by shifting the mantissa bits and an adder tree configured to output an accumulation value of the primitive products in an integer format by adding the shifted mantissa bits.
Owner:OPENAI OPCO LLC

Internal memory, chip and related electronic device

Disclosed in the embodiments of the present application are an internal memory, a chip and a related electronic device. A target processing unit (202) of an internal memory (102) comprises a plurality of groups of primary multipliers, a plurality of adder trees and a plurality of secondary multipliers, wherein one group of primary multipliers is connected to one secondary multiplier by means of one adder tree; one group of primary multipliers comprises a plurality of pairs of first multipliers; each adder tree comprises a first adder, a second adder and a third adder; each pair of first multipliers separately receives data to perform multiplication operations, and separately inputs output results into one first adder for addition operations; each second adder receives output results of a plurality of first adders or the upper-level second adder, so as to perform addition operations, and separately inputs the output results into the lower-level second adder or one third adder; and each third adder receives output results of a plurality of second adders to perform addition operations, and inputs the output results into one secondary multiplier for multiplication operations. Implementing the embodiments of the present application can improve the processing capability of the internal memory.
Owner:HUAWEI TECH CO LTD

Alignment in hardware accelerators

Systems, apparatuses, and methods are disclosed for improved matrix-vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values in floating point formats. The CIM macro has a functional block configured to align mantissa bits of primitive products between the activation values and the weight values by shifting the mantissa bits and an adder tree configured to output an accumulation value of the primitive products in an integer format by adding the shifted mantissa bits.
Owner:OPENAI OPCO LLC

Neural network calculation acceleration method and system

The invention provides a neural network calculation acceleration method and system, and the method comprises the steps: extracting an amplitude bit of input vector data and an amplitude bit of weight vector data, and generating a zero tag through logic operation; the zero label represents whether the product of the input vector data and the weight vector data is zero or not; dynamically aggregating the non-zero products according to the zero label to form an aggregated product; and dynamically adjusting the calculation scale of the adder tree according to the number of non-zero products in the aggregated products, and only performing accumulation operation on the non-zero products. According to the method, invalid calculation can be effectively reduced, the calculation efficiency and the utilization efficiency of hardware resources are improved, and the method is suitable for accelerated calculation of various neural network models.
Owner:PEKING UNIV

Method and apparatus for performing deep learning operations

ActiveCN114595811BAccumulator (computing)Binary multiplier
Methods and apparatuses for performing deep learning operations are provided. The computing apparatus includes an adder-tree-based tensor kernel and a multiplier-accumulator (MAC)-based vector kernel. The adder-tree-based tensor kernel is configured to perform tensor operations, and the multiplier-accumulator (MAC)-based vector kernel is configured to perform vector operations using the output of the tensor kernel as input.
Owner:SAMSUNG ELECTRONICS CO LTD

In-memory calculation matrix multiplication acceleration system for fine-grained structured sparseness

The invention belongs to the technical field of in-memory computing, and discloses a fine-grained structured sparse-oriented in-memory computing matrix multiplication acceleration system, which comprises an in-memory computing array group and an add tree, the in-memory computing array group consists of M rows and N columns and consists of a plurality of in-memory computing arrays, and each in-memory computing array comprises a storage module mem, a computing unit and a ping-pong pulsation input transmission chain PPSIC, the PPSIC is used for transmitting input data to the calculation module cmp and transmitting the input data to the PPSIC of the next column; and the computing units of the in-memory computing arrays on the same column are connected with the same add tree. According to the technical scheme, the internal structure of the digital in-memory computing array is optimized, the flexibility of the in-memory computing array is improved, and N: M fine-grained structured sparsity can be efficiently supported.
Owner:PEKING UNIV

Hybrid precision MAC tree structure for maximizing memory bandwidth usage to accelerate operation of generative large-scale language models

The present invention provides a hybrid precision MAC (multiple-and-multiple-accumulation) tree structure for maximizing memory bandwidth usage to accelerate the operation of a generative large scale language model. The present invention provides a hybrid precision MAC (multiple-and-multiple-accumulation) tree structure for maximizing memory bandwidth usage to accelerate the operation of the generative large scale language model. A MAC tree-based arithmetic unit according to one embodiment may include: a plurality of floating point multipliers connected in parallel and processing multiplication of data transferred from an external memory; a plurality of first converters for converting the output of each of the plurality of floating point multipliers from a floating point to a fixed point; a fixed-point adder tree which is connected to the plurality of first converters and processes the addition of the multiplication results of the plurality of floating-point multipliers; a fixed-point accumulator that accumulates the output of the fixed-point adder tree; and a second converter that converts the output of the fixed-point accumulator from a fixed point to a floating point.
Owner:超速有限公司

Accelerate neural networks with compression at different levels

A neural network accelerator includes 2n multiplier circuits, 2n shifter circuits and an adder tree circuit. Each respective multiplier circuit multiplies a first value by a second value to output a first product value. Each respective first value is represented by a first predetermined number of bits beginning at a most significant bit of the first value having a value equal to 1. Each respective second value is represented by a second predetermined number of bits, and each respective first product value is represented by a third predetermined number of bits. Each respective shifter circuit receives the first product value of a corresponding multiplier circuit and left shifts the corresponding product value by the first predetermined number of bits to form a respective second product value. The adder circuit adds each respective second product value to form a partial-sum value represented by a fourth predetermined number of bits.
Owner:SAMSUNG ELECTRONICS CO LTD

Multiplier unit and method and apparatus for calculating the dot product of floating point values

A multiplier unit and a method and apparatus for calculating a dot product of floating point values are disclosed. The apparatus includes an array of multiplier units each including integer logic to multiply integer values of corresponding elements of two vectors, exponent logic to add exponent values of corresponding elements of the two vectors to form un-biased exponent values, and a local shifter to form a first shifted value by shifting an integer product value in a predetermined direction by a number of bits based on a difference between the un-biased exponent value corresponding to the integer product value and a maximum un-biased exponent value of the array of multiplier units. An adder tree adds the shifted values output from the local shifters of the array of multiplier units to form an output, and an accumulator accumulates the output of the adder unit.
Owner:SAMSUNG ELECTRONICS CO LTD

Adder tree circuit and floating point calculation method

The invention discloses an adder tree circuit and a floating point calculation method, and relates to the technical field of chips. The adder tree circuit comprises an analysis circuit used for determining a plurality of floating point input data corresponding to a function operator and analyzing the plurality of floating point input data to obtain indexes and mantissas corresponding to the plurality of floating point input data; wherein the function operator is used for executing additive operation on a plurality of floating point input data; the order matching processing circuit is used for performing order matching on the indexes corresponding to the floating point input data based on the indexes and mantissas corresponding to the floating point input data to obtain mantissas after order matching corresponding to the floating point input data; the additive operation circuit is used for performing additive operation on the mantissa after order based on the plurality of floating point input data to obtain a first operation result; and the normalization processing circuit is used for performing normalization processing on the first operation result to obtain a second operation result. The area overhead of the logic operation circuit can be reduced, and the accuracy of the operation result is ensured.
Owner:NINGBO HORIZON SATENG TECHNOLOGY CO LTD

SRAM-based in-memory computing method and apparatus in digital domain

The application provides an SRAM-based digital domain in-memory computing method and device. First, input data and bit high-low order of the input data are acquired, and SRAM in-memory computing arrays corresponding to the input data respectively are determined according to the bit high-low order of the input data. Different preset approximate adder trees are included in different SRAM in-memory computing arrays, the different preset approximate adder trees include preset approximate full adders with different calculation accuracies, and the structure of the adder tree is a Wallace tree structure. Finally, the input data are subjected to approximate operation through the preset approximate adder tree, and a data approximate operation result is obtained. Through the above method, in combination with the application of the preset approximate full adder and the Wallace tree structure, the preset approximate adder tree with different accuracies can be selected to perform data calculation according to the bit high-low order of the input data while reducing the area and power consumption of the adder tree, and the calculation accuracy is improved.
Owner:UNIV OF SCI & TECH OF CHINA

Hierarchical adder tree structure module with fine-grained exact reconfigurable approximation computation

ActiveCN115220689BRealize approximate addition calculationSolve the problem of not being able to cope with data vector calculationsTheoretical computer scienceApproximate computing
The application relates to a hierarchical addition tree structure module with fine-grained accurate reconfigurable approximate calculation, comprising a multi-data input module, an adder tree structure generation module, a calculation and result output module; the multi-data input module receives addends and defines approximate adder calculation, and receives user required precision configuration; the adder tree structure generation module input end receives addition numbers; after initialization, the number of addition tree layers will be transmitted to the approximate adder tree structure module to complete generation of the fine-grained accurate reconfigurable approximate adder module; the calculation and result output module performs approximate addition operation on the approximate adder generated in the previous stage to complete the final approximate calculation task, and controls the precision required by each layer in the approximate calculation process. The application realizes approximate addition calculation on a data vector, higher energy efficiency and appropriate precision, and solves the problem that the existing approximate adder configuration scheme cannot cope with data vector calculation.
Owner:NANJING RES INST OF ELECTRONICS TECH

Deconvolution calculation method, calculation apparatus and calculation system, and storage medium

A deconvolution calculation method, calculation apparatus (100) and calculation system (300), and a storage medium (400). The deconvolution calculation method comprises: by means of a systolic array (10), calculating each weight of each row of a convolution kernel with a first feature matrix, so as to obtain a second feature matrix corresponding to each weight of each row of the convolution kernel; and by means of an adder tree (20), calculating a plurality of second feature matrices corresponding to a plurality of weights of a plurality of rows of the convolution kernel, so as to obtain a third feature matrix.
Owner:BYD CO LTD

Calculation circuit, memory device including the calculation circuit, and calculation method

A calculation circuit, a memory device including the calculation circuit, and a calculation method are provided. The calculation circuit comprises an input allocator receiving and dividing n-bit input data (where n is a natural number equal to or greater than 2) into a plurality of operation elements based on a data type of the input data, an adder tree performing a multiplication operation between the operation elements, and an accumulator generating a first output value by adding an output value of the adder tree to a value stored in an accumulation register, wherein the first output value includes a sign bit and data bits, and the accumulator includes a first lightweight normalizer that performs bit shifting on the first output value by comparing a value of the sign bit with values of m bits (where m is a natural number) among the data bits.
Owner:SAMSUNG ELECTRONICS CO LTD

Method and apparatus for performing deep learning operations

A method and apparatus for performing deep learning operations. A computation apparatus includes an adder tree-based tensor core configured to perform a tensor operation, and a multiplier and accumulator (MAC)-based vector core configured to perform a vector operation using an output of the tensor core as an input.
Owner:SAMSUNG ELECTRONICS CO LTD

Variable bit-width adder tree generation system based on multiple types of approximate computing units

The application relates to a variable bit width adder tree generation system based on multiple types of approximate calculation units, which comprises a signal input module, an adder tree construction module and a calculation and result output module. The input end A of the signal input module receives two groups of addends and defines the bit width; the input end B receives the required precision bit number and initializes the precision comparison module. The input end of the adder tree construction module receives two groups of addends; the approximate calculation unit library and the Boolean gate logic unit library are initialized and called respectively; in the iteration process, different types of approximate calculation units are selected in different levels according to the precision requirement input by the user. The calculation and result output module performs approximate addition operation on the approximate adder tree module generated in the last stage, and finally completes the approximate calculation task. The application realizes dynamic configuration of the bit width of the approximate addition operation, makes the selection of the precision and the power consumption more flexible, and solves the technical problem of poor precision of the existing approximate addition dynamic configuration scheme.
Owner:NANJING RES INST OF ELECTRONICS TECH

A high-bandwidth utilization sparse matrix vector multiplication acceleration device

The application provides a high-bandwidth utilization rate sparse matrix vector multiplication acceleration device, which comprises a decoder, a read conflict-free input vector buffer, a calculation unit array, a write conflict-free adder tree, a ping-pong supporting accumulator group, a storage part and a result vector buffer; the decoder is used for decoding a preprocessed matrix; the decoder transmits vector elements in the matrix into the read conflict-free input vector buffer after decoding; non-zero elements in a target matrix are decoded and transmitted into the calculation unit array; the calculation unit array is used for reading corresponding vector elements from the read conflict-free input vector buffer according to column numbers of the non-zero elements, multiplying the vector elements with the non-zero data, and transmitting the multiplication results and row numbers of the non-zero elements into the write conflict-free adder tree; the write conflict-free adder tree is used for adding multiplication results with the same row numbers, and transmitting the addition results into accumulators; and the ping-pong supporting accumulator group is used for accumulating the addition results.
Owner:CHONGQING UNIV

Adder tree circuit, adder circuit and method of operating a full adder

In some aspects of the application, adder tree circuits are disclosed. In some aspects, the adder tree circuit includes a plurality of full adders (FAs) including: a first subset of FAs, wherein each FA of the first subset of full adders includes a first number of transistors; and a second subset of FAs, wherein each FA of the second subset of full adders includes a second number of transistors, the first number being greater than the second number; wherein each FA of the first subset receives a first input from a first one of the second subset of FAs and a second input from a second one of the second subset of FAs, and each FA provides a first output to a third one of the second subset of FAs and a second output to a fourth one of the second subset of FAs. Embodiments of the present application also relate to adder circuits and methods of operating full adders.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Convolution Circuit, Convolution Computing Method, Chip, and Electronic Device

A convolution circuit includes a plurality of multipliers, a first adder coupled to the plurality of multipliers, and a second adder coupled to the first adder. Each multiplier includes a plurality of precoders, a plurality of encoder groups, and an adder tree circuit. Each precoder is in a one-to-one correspondence with one encoder group. Output ends of the plurality of encoder groups and input lines of the adder tree circuit are of a same quantity and in a one-to-one correspondence. In addition, the adder tree circuit is coupled to the first adder. The second adder is further coupled to a memory. A partial product that is related only to a weight parameter may be first accumulated with a constant 1 in the multiplier, and then added to results output by adder tree circuits in the second adder.
Owner:HUAWEI TECH CO LTD

Multi-precision MAC tree type processing unit and systolic array structure

The invention relates to the technical field of digital integrated circuits, and particularly provides a multi-precision MAC tree type processing unit and a systolic array structure.The processing unit comprises an MAC tree type structure, and the MAC tree type structure comprises a multi-precision multiplication module used for carrying out data multiplication operation on multiple sets of first precision data and second precision data, determining a corresponding first product result and a second product result; the adder tree module comprises a first hierarchical structure and a second hierarchical structure used for determining a product accumulation result; the first hierarchical structure comprises N node structures, each node structure comprises an adder and a first data selector, the adder is used for receiving the first product result, performing additive operation and generating a first addition result, and two paths of input signals of the first data selector are the first addition result and the second product result respectively. The problems of operation delay and storage bottleneck in the matrix operation process can be solved.
Owner:BEIJING ZHONGKE YIHAI MICROELECTRONICS TECHNOLOGY RESEARCH INSTITUTE CO LTD

Depthwise-convolution implementation on a neural processing core

A core of neural processing units is configured to efficiently process a depthwise convolution by maximizing spatial feature-map locality using adder trees. Data paths of activations and weights are inverted, and 2-to-1 multiplexers are every 2 / 9 multipliers along a row of multipliers. During a depthwise convolution operation, the core is operated using a RS×HW dataflow to maximize the locality of feature maps. For a normal convolution operation, the data paths of activations and weights may be configured for a normal convolution configuration and in which multiplexers are idle.
Owner:SAMSUNG ELECTRONICS CO LTD

Adder tree architecture supporting DNN sparse perception and implementation method thereof

The invention relates to the technical field of artificial intelligence chips, and particularly discloses an adder tree architecture supporting DNN sparse perception and an implementation method thereof. The architecture comprises an index and data reordering module, a reconfigurable sparse addition array, an address mapping module, a partial storage area and an interface control unit of the partial storage area. And by carrying out index reordering on sparse multiplication results and adopting an adder array which is only activated when indexes are consistent, dynamic accumulation of random sparse data is realized. According to the architecture, a local accumulator in a traditional processing unit is removed, parallel operation of calculation and accumulation decoupling is supported, and the hardware utilization rate and the accumulation efficiency are effectively improved. The method is suitable for efficient addition and accumulation of various neural network layers such as convolution, full connection, normalization and the like, and has the advantages of high structural universality, high energy efficiency ratio, low area overhead and the like.
Owner:XI AN JIAOTONG UNIV

Circuit of an adder and method of operation thereof

A circuit of an adder and a method of operation thereof are provided. The circuit includes a plurality of first adder tree structures and a plurality of second adder tree structures. The plurality of first adder tree structures perform additions of bits of a plurality of digits that are below a bit value. Adders in the plurality of first adder tree structures produce full swing outputs. The plurality of second adder tree structures are coupled to the plurality of first adder tree structures. The plurality of second adder tree structures perform additions of bits of the plurality of digits that are at or above the bit value. The plurality of second adder tree structures include fewer transistors than the plurality of first adder tree structures. One of the adders in the plurality of second adder tree structures produces a non-full swing output in response to a bit input to the adder in the plurality of second adder tree structures matching a pattern.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Calculation circuit, memory cell including calculation circuit, and calculation method

The invention provides a calculation circuit, a storage device including the calculation circuit, and a calculation method. The computing circuit includes: an input distributor receiving n bits of input data and dividing the input data into a plurality of operational elements based on a data type of the input data, where n is a natural number equal to or greater than 2; an adder tree that performs a multiplication operation between the operation elements; and an accumulator generating a first output value by adding an output value of the adder tree to the value stored in the accumulation register, where the first output value includes sign bits and data bits, and the accumulator includes a first lightweight normalizer configured to normalize the sign bits and the data bits, and a second lightweight normalizer configured to normalize the sign bits and the data bits. And a shift unit that performs a shift on the first output value by comparing a value of the sign bit with a value of m bits among the data bits, where m is a natural number greater than n, less than n, or equal to n.
Owner:SAMSUNG ELECTRONICS CO LTD

Novel hybrid precision convolution multiplier-adder

The invention discloses a novel hybrid precision convolution multiplier-adder, and belongs to the technical field of artificial intelligence chip design and deep learning acceleration. The multiplier and adder supports two core data formats of INT16 and FP16, and efficient sharing of integer and floating point multiplication and addition resources is realized by improving a floating point number operation method. The core of the method is to deform a traditional FP16 floating-point number representation method, so that a 16-bit * 16-bit integer multiplier can be used in the calculation process of the method, and meanwhile, a multi-stage calculation architecture (including modules such as MTS calculation, product and maximum order code solution, symbol processing and shifting, adder tree summation and the like) is adopted, so that delay caused by multiple alignment of order codes is reduced. According to the design, multiplier and adder tree resources are multiplexed, on the premise that the accuracy loss is controllable, hardware resource redundancy and operation complexity are reduced, the advantage of comprehensive performance is more remarkable along with increase of the order number of the multiplier and adder, and the method is suitable for efficient acceleration scenes of deep learning convolution operation.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Deconvolution calculation method, calculation device, calculation system and storage medium

The invention discloses a deconvolution calculation method, a calculation device, a calculation system and a computer readable storage medium. The deconvolution calculation method is applied to the deconvolution calculation device, the deconvolution calculation device comprises a systolic array and an adder tree, the systolic array is connected with the adder tree, and the deconvolution calculation method comprises the following steps: inputting a convolution kernel and a first feature matrix into the systolic array; calculating each weight of each row of the convolution kernel and the first feature matrix through a systolic array to obtain a second feature matrix corresponding to each weight of each row of the convolution kernel; and calculating a plurality of second feature matrixes corresponding to the plurality of weights of the plurality of rows of the convolution kernel through the adder tree to obtain a third feature matrix, the third feature matrix being a deconvolution calculation result of the first feature matrix. Therefore, the computing resources of the deconvolution computing device are used for computing the effective data, and the storage resources are not wasted in storing the zero values which do not contribute to the computing, so that the hardware computing efficiency of the deconvolution computing device is effectively improved.
Owner:BYD SEMICON CO LTD

Reconfigurable Processor Circuit Architecture

A representative reconfigurable processing circuit and a reconfigurable arithmetic circuit are disclosed, each of which may include input reordering queues; a multiplier shifter and combiner network coupled to the input reordering queues; an accumulator circuit; and a control logic circuit, along with a processor and various interconnection networks. A representative reconfigurable arithmetic circuit has a plurality of operating modes, such as floating point and integer arithmetic modes, logical manipulation modes, Boolean logic, shift, rotate, conditional operations, and format conversion, and is configurable for a wide variety of multiplication modes. Dedicated routing connecting multiplier adder trees allows multiple reconfigurable arithmetic circuits to be reconfigurably combined, in pair or quad configurations, for larger adders, complex multiplies and general sum of products use, for example.
Owner:CORNAMI INC

Multiplication and accumulation(MAC) operator and processing-in-memory (PIM) device including the mac operator

A multiplication-accumulation (MAC) includes a multiplication circuit, a pre-processing circuit, and an adder tree. The multiplication circuit performs a multiplication operation on a plurality of weight data and a plurality of vector data each having a floating-point format to output a plurality of multiplication data. The pre-processing circuit performs shifting on mantissa data of the plurality of multiplication data by a difference between first maximum exponent data having a greatest value among the exponent data of the plurality of multiplication data and the remaining exponent data to output a plurality of pre-processed mantissa data. The adder tree adds the plurality of mantissa data to output mantissa addition bits.
Owner:SK HYNIX INC