Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

28 results about "Floating point multiplication" patented technology

Low-bit-width high-energy-efficiency floating point storage and calculation integrated circuit based on partial pre-alignment architecture

The invention belongs to the technical field of storage and calculation integration, and particularly relates to a low-bit-width and high-energy-efficiency floating point storage and calculation integrated circuit based on a partial pre-alignment framework. The circuit comprises a memory array, a pre-calculation unit, an adder tree, a configurable arithmetic unit and a normalization unit, and supports mixed precision operation of FP8MACFP4 and FP8MACFP8. The method is characterized in that a partial pre-alignment strategy dominated by an activation value is adopted, the maximum index of the activation value is dynamically counted, the mantissa of the maximum index is aligned, and multiple partial pre-alignment intermediate results are pre-calculated and latched for reuse; in combination with a customized lookup table and a multiplexer, a pre-calculation result is directly selected to replace real-time multiplication and displacement; and through the reconfigurable hardware, the FP8MACFP8 high-precision operation is realized by utilizing the FP8MACFP4 unit combination. According to the method, complete online floating point multiplication and addition operation is realized, and excellent energy efficiency ratio and operation speed are obtained while high precision is kept.
Owner:FUDAN UNIVERSITY

Approximate floating point multiplier, chip and computing device

PendingCN120104094ADigital data processing detailsBinary multiplierAnd logic unit
The invention discloses an approximate floating point multiplier, a chip and computing equipment, an approximate mantissa multiplier of the approximate floating point multiplier comprises an AND logic unit and a compressor unit, and the AND logic unit is used for performing AND operation on two input operands bit by bit to generate a partial product array with the size of 11 rows and 21 columns; the compressor unit is used for compressing the 11th column to the 21st column by column to obtain a final approximate mantissa, the compressor unit comprises two novel approximate 4-2 compressors ignoring carry design, and the error rate of the approximate 4-2 compressors is within an acceptable range by utilizing mutual compensation inside the compressors. The invention aims to excavate and use the characteristics of floating point multiplication to further improve the energy efficiency of floating point multiplication, realize the optimization of the approximate floating point multiplier in the overhead aspects of precision, power consumption, area and the like, and solve the problems of relatively complex circuit and low compression efficiency of the traditional approximate 4-2 compressor.
Owner:NAT UNIV OF DEFENSE TECH

Systems and methods for energy-efficient, bit-parallel, multiply-accumulate for artificial intelligence and deep neural networks

A system and method for providing a tunable floating-point multiply-accumulate (MAC) unit are disclosed. The unit maintains full arithmetic precision while enabling dynamic elimination of ineffectual computation through operand decomposition and selective activation of partial product generation logic. The disclosed MAC unit is suitable for drop-in replacement in existing deep-learning accelerators and improves energy efficiency without requiring architectural changes.
Owner:KAXIRAS STEFANOS +3

Fused multiply add operations

A data processing apparatus is provided that performs a fused multiply add operation. Multiplication circuitry multiplies pairs of floating-point multiplication values together to produce unrounded products. Addition circuitry adds a sum of the products to an accumulation value to produce a rounded floating-point result and control circuitry, responsive to a fused multiply add instruction, controls the multiplication circuitry and the addition circuitry to perform the fused multiply add operation. The addition circuitry accesses the accumulation value at a later processing cycle than the multiplication circuitry accesses the pairs of floating-point multiplication values.
Owner:ARM LTD

Floating-point number multiplier, calculation chip and floating-point number multiplication operation method

The invention relates to a floating-point number multiplier, a computing chip and a floating-point number multiplication operation method, and belongs to the technical field of computers. According to the floating-point number multiplier, floating-point multiplication of a target floating-point number type is performed on input data by utilizing multiplication mode indication, the floating-point number type is specified as the floating-point number type of floating-point numbers in the input data, the floating-point numbers in the input data are divided into data groups by preprocessing the input data, and the floating-point numbers in the input data are divided into the data groups based on a target multiplication mode. According to the invention, the floating point data in the same data set is subjected to floating point number type floating point number multiplication operation, so that the floating point number multiplier can perform floating point number multiplication operation of various floating point number types by specifying different multiplication modes for input data of different floating point number types; therefore, the operation requirements of floating-point number multiplication of various floating-point number types can be met.
Owner:北京凌川科技有限公司 +1

Hybrid precision MAC tree structure for maximizing memory bandwidth usage to accelerate operation of generative large-scale language models

The present invention provides a hybrid precision MAC (multiple-and-multiple-accumulation) tree structure for maximizing memory bandwidth usage to accelerate the operation of a generative large scale language model. The present invention provides a hybrid precision MAC (multiple-and-multiple-accumulation) tree structure for maximizing memory bandwidth usage to accelerate the operation of the generative large scale language model. A MAC tree-based arithmetic unit according to one embodiment may include: a plurality of floating point multipliers connected in parallel and processing multiplication of data transferred from an external memory; a plurality of first converters for converting the output of each of the plurality of floating point multipliers from a floating point to a fixed point; a fixed-point adder tree which is connected to the plurality of first converters and processes the addition of the multiplication results of the plurality of floating-point multipliers; a fixed-point accumulator that accumulates the output of the fixed-point adder tree; and a second converter that converts the output of the fixed-point accumulator from a fixed point to a floating point.
Owner:超速有限公司

Floating point multiplications

Systems, apparatuses, and methods are disclosed for improved matrix-vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values, and a mode decoding unit is configured to provide a mode of the VMM operation according to a first floating point format of the activation values and a second floating point format of the weight values. A column cell of the CIM macro can be configured to output a product between an activation value and a weight value using a half adder and multiple multiplexers that provide selections to a full adder based on control signals.
Owner:OPENAI OPCO LLC

Floating point multiplication and addition instruction fusion method and device, equipment and storage medium

The invention relates to the technical field of computers, and provides a floating point multiplication and addition instruction fusion method and device, equipment and a storage medium. According to the method, a fusion scaling factor and a fusion offset of two continuous floating point multiplication and addition operations are calculated, the two continuous floating point multiplication and addition operations are merged based on the fusion scaling factor and the fusion offset, and a fusion instruction is output. And two FMA operations are combined into a single instruction by utilizing mathematical equivalent transformation, so that the problems of relatively large computing load, relatively high computing delay and low hardware utilization rate of a GPU (Graphic Processing Unit) caused by relatively large calling of floating point multiply-add instructions in exponential approximate calculation in existing softmax are solved.
Owner:GUANGZHOU WERIDE TECH LTD CO

An FPGA-based floating-point multiplier, calculation method and device

This application relates to the field of floating-point multiplication, and specifically discloses a floating-point multiplier, a calculation method, and a device based on an FPGA. The floating-point multiplier includes: a sign calculation module for determining the sign of the target output floating-point number by means of an exclusive-OR gate and the sign bits of the first input floating-point number and the second input floating-point number; an exponent addition module for adding the exponents of the first input floating-point number and the second input floating-point number and subtracting the bias value of the corresponding format to obtain the exponent output of the target output floating-point number; a binary multiplier module that performs a multiplication operation based on the mantissa bit widths of the first input floating-point number and the second input floating-point number using the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm to obtain the mantissa product of the target output floating-point number; and a result normalization module for performing a normalization operation based on the mantissa product. It can reduce the calculation delay and also reduce the percentage increase in the hardware area.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Hardware circuit, method, device and medium for floating point matrix multiplication

The invention is suitable for the field of chip and digital circuit design, and provides a hardware circuit, a method, equipment and a medium for floating point matrix multiplication, and the hardware circuit comprises a floating point multiplication unit which is used for carrying out multiplication on an obtained multiplier in a first standard floating point format and an obtained multiplicand in a second standard floating point format, outputting a product expressed in a first user-defined floating point format; the floating point accumulation unit is used for adding the product and an accumulation input value and outputting an accumulation sum expressed by adopting a second user-defined floating point format; and the format conversion unit is used for converting the accumulated sum of the second custom floating point format into a third standard floating point format and outputting the third standard floating point format. According to the scheme, balance can be achieved between keeping of calculation precision and limiting of hardware overhead.
Owner:GUANGDONG LEAPFIVE TECH CO LTD

Method and apparatus for implied bit handling in floating point multiplication

Devices and methods are provided for performing, by a processor in response to a floating point multiply instruction, multiplication of floating point numbers. An example processor includes first, second, third, and fourth computational paths. In operation, the first determines values of implied bits of mantissas of floating point numbers and generates first partial product terms, the second multiplies remainders of the mantissas to generate second partial product terms, the third detects a number of leading zeros in the mantissas and determines a shift amount for each of the mantissas, and the fourth calculates exponents for a flush-to-zero mode.
Owner:TEXAS INSTRUMENTS INC

Floating point multiplications

Systems, apparatuses, and methods are disclosed for improved matrix–vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values, and a mode decoding unit is configured to provide a mode of the VMM operation according to a first floating point format of the activation values and a second floating point format of the weight values. A column cell of the CIM macro can be configured to output a product between an activation value and a weight value using a half adder and multiple multiplexers that provide selections to a full adder based on control signals.
Owner:OPENAI OPCO LLC

Approximate floating-point multiplication and accumulation circuit and method

The present application discloses an approximate floating-point multiplication and accumulation circuit and method, which belongs to the field of computer architecture and digital circuit design technology, including: an approximate floating-point multiplier precision arbitration module, which receives a preprocessed floating-point number and outputs a first precision control signal; an approximate floating-point multiplication module, which adopts an approximate addition scheme with different mantissa multiplications according to the first precision control signal and outputs an approximate floating-point multiplication result; an approximate floating-point adder precision arbitration module, which receives an approximate floating-point multiplication result and a partial sum result of the previous time and outputs a second precision control signal; an approximate floating-point addition module, which adopts an approximate addition scheme with different mantissa additions according to the second precision control signal and outputs a partial sum result. The present application dynamically adjusts the calculation process according to the characteristics of the input data. When the input data differs greatly, the circuit adopts precise calculation to ensure the accuracy of the calculation result; when the input data differs slightly, the circuit switches to an approximate calculation mode to simplify the calculation process.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Tensor arithmetic unit based on RISC-V instruction set and intelligent processor

The invention provides a tensor arithmetic unit based on an RISC-V instruction set, and the tensor arithmetic unit comprises a microinstruction splitting and scheduling unit which receives a decoded macroscopic tensor instruction and an operation size parameter thereof, splits the macroscopic tensor instruction into microinstruction sequences according to a fixed physical scale of a reconfigurable calculation array, and transmits the microinstruction sequences to the reconfigurable calculation array; a zigzag traversal sequence and microinstruction scheduling are achieved through quintuple hardware circulation, and the loading, using and replacing sequence of the data blocks in the block register array is planned; the reconfigurable computing array is used for executing matrix multiply-accumulate operation of various data formats by designing a reconfigurable data path and fusing and multiplexing a floating point multiplier and a multi-precision accumulation tree under data paths with different bit widths; and the hardware multi-buffer unit is a multi-buffer architecture automatically managed by hardware and is matched with a zigzag traversal sequence to realize pipeline overlapping of data prefetching and calculation execution. The invention further provides an intelligent processor. Therefore, the invention can efficiently execute tensor operation with variable precision and variable scale.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Converting floating point addition to integer addition for artificial intelligence applications

The invention relates to converting floating point addition to integer addition for artificial intelligence applications. One embodiment provides a graphics processor, comprising: a base die comprising a plurality of core grain sockets; and a plurality of core particles coupled with the plurality of core particle sockets. At least one of the plurality of cores includes a matrix accelerator having circuitry for performing a floating point operation, the floating point operation including a plurality of floating point multiplications. The circuit comprises: an input circuit for storing a floating point input value; a first adder for generating an intermediate exponent sum; the multiplier is used for generating an intermediate mantissa product; and an intermediate accumulator configured to accumulate the plurality of intermediate mantissa products as integer values in a plurality of rows of the bitwise storage device, the plurality of rows being respectively associated with different exponent sums.
Owner:INTEL CORP

Mixed-precision floating-point operation circuit in a dedicated processing block

Mixed precision floating point operation circuits in special purpose processing blocks are disclosed. Embodiments relate to integrated circuits having circuits that efficiently perform mixed precision floating point arithmetic operations. Such circuits can be implemented in special purpose processing blocks. The special purpose processing blocks can include configurable interconnect circuits to support a variety of different usage modes. For example, the special purpose processing blocks can implement fixed point addition, floating point addition, fixed point multiplication, floating point multiplication, sum of two multiplications, with or without conversion to a second floating point precision, followed by a subsequent addition in the second floating point precision if needed, to name a few. In some embodiments, two or more special purpose processing blocks can be arranged in a cascade chain and together perform more complex operations such as a recursive mode dot product of two vectors of floating point numbers having a first floating point precision and output the dot product in a second floating point precision.
Owner:ALTERA CORP

Transforming floating-point addition into integer addition for artificial intelligence applications

One embodiment provides a graphics processor comprising a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets coupled with the plurality of chiplet sockets. At least one of the plurality of chiplets include a matrix accelerator having circuitry to perform a floating-point operation including a plurality of floating-point multiplications. The circuitry includes input circuitry to store floating-point input values, a first adder to generate an intermediate exponent sum, a multiplier to generate an intermediate mantissa product, and an intermediate accumulator configured to accumulate a plurality of intermediate mantissa products as integer values within a plurality of rows of bitwise storage, the plurality of rows respectively associated with different exponent sums.
Owner:INTEL CORP

Floating point multiplication method and floating point multiplication circuit

The application provides a floating-point multiplication method and a floating-point multiplication circuit, and relates to the technical field of processors. The method comprises the following steps: obtaining a shift result of the product of a first floating-point number and a second floating-point number; the shift result is the result of a shift operation on the decimal part of the product; performing a mantissa rounding operation on the shift result to obtain a mantissa rounding result of the shift result; obtaining a first judgment result of whether the mantissa satisfies a first preset rounding condition in the process of the mantissa rounding operation; if the first judgment result is that the mantissa satisfies the first preset rounding condition, obtaining a target mantissa of the multiplication result of the first floating-point number and the second floating-point number according to the mantissa rounding result. The floating-point multiplication method has the advantage of short operation period.
Owner:BEIJING INSTITUTE OF OPEN SOURCE CHIP

Distributed double-precision floating-point multiplication

The present embodiments relate to circuitry that efficiently performs double-precision floating-point multiplication operations, single-precision floating-point multiplication operations, and fixed-point multiplication operations. Such circuitry may be implemented in specialized processing blocks. If desired, each specialized processing block efficiently may perform a single-precision floating-point multiplication operation, and multiple specialized processing blocks may be coupled together to perform a double-precision floating-point multiplication operation. Inter-block signaling circuits may generate rounding information and propagate the rounding information together with partial product results from a current specialized processing block to another specialized processing block.
Owner:ALTERA CORP

Transcendental function computation system and method based on interpolation approximation, and chip and terminal device

Provided in the present invention are a transcendental function computation system and method based on interpolation approximation, and a chip and a terminal device. The method comprises: an input module inputting floating-point numbers having a preset number of bits; a single-precision floating-point number computation module performing transcendental function computation on the input floating-point numbers in the computation mode of single-precision floating-point multiplication, and outputting computation results in a half-precision floating-point format; a half-precision floating-point number computation module performing transcendental function computation on the input floating-point numbers in the computation mode of half-precision floating-point multiplication, and outputting computation results in a half-precision floating-point format; and an output module outputting the computation results. By means of a single-precision floating-point number computation module and a half-precision floating-point number computation module, high-precision and high-performance operations can be performed on single-precision floating-point numbers and half-precision floating-point numbers, such that the requirements of processor chips and their interface APIs are met, thereby enabling the transcendental function computation method to not only reduce costs while ensuring precision, but also be applicable to various types of transcendental functions.
Owner:VERISILICON MICROELECTRONICS (SHANGHAI) CO LTD +4

Method and apparatus for implied bit handling in floating point multiplication

Devices and methods are provided for performing, by a processor in response to a floating point multiply instruction, multiplication of floating point numbers. In an example, a device includes a processor that includes a multiply circuit. The multiply circuit is configured to multiply floating point numbers in response to a floating point multiply instruction, and is further configured to determine values of implied bits of mantissas of the floating point numbers, and multiply the mantissas in parallel with the determining operation.
Owner:TEXAS INSTRUMENTS INC

Novel hybrid precision convolution multiplier-adder

The invention discloses a novel hybrid precision convolution multiplier-adder, and belongs to the technical field of artificial intelligence chip design and deep learning acceleration. The multiplier and adder supports two core data formats of INT16 and FP16, and efficient sharing of integer and floating point multiplication and addition resources is realized by improving a floating point number operation method. The core of the method is to deform a traditional FP16 floating-point number representation method, so that a 16-bit * 16-bit integer multiplier can be used in the calculation process of the method, and meanwhile, a multi-stage calculation architecture (including modules such as MTS calculation, product and maximum order code solution, symbol processing and shifting, adder tree summation and the like) is adopted, so that delay caused by multiple alignment of order codes is reduced. According to the design, multiplier and adder tree resources are multiplexed, on the premise that the accuracy loss is controllable, hardware resource redundancy and operation complexity are reduced, the advantage of comprehensive performance is more remarkable along with increase of the order number of the multiplier and adder, and the method is suitable for efficient acceleration scenes of deep learning convolution operation.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Hardware circuit, method, device and medium for floating point matrix multiplication operation

The application is suitable for the field of chip and digital circuit design, and provides a hardware circuit, method, equipment and medium for floating point matrix multiplication, wherein the hardware circuit comprises: a floating point multiplication unit, which is used for multiplying a first standard floating point format multiplier and a second standard floating point format multiplicand, and outputs a product in a first self-defined floating point format; a floating point accumulation unit, which is used for adding the product and an accumulation input value, and outputs an accumulation sum in a second self-defined floating point format; and a format conversion unit, which is used for converting the accumulation sum in the second self-defined floating point format into a third standard floating point format and then outputting. The scheme can balance between maintaining calculation accuracy and limiting hardware overhead.
Owner:GUANGDONG LEAPFIVE TECH CO LTD