Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

22 results about "Floating point multiplication" patented technology

Low-bit-width high-energy-efficiency floating point storage and calculation integrated circuit based on partial pre-alignment architecture

The invention belongs to the technical field of storage and calculation integration, and particularly relates to a low-bit-width and high-energy-efficiency floating point storage and calculation integrated circuit based on a partial pre-alignment framework. The circuit comprises a memory array, a pre-calculation unit, an adder tree, a configurable arithmetic unit and a normalization unit, and supports mixed precision operation of FP8MACFP4 and FP8MACFP8. The method is characterized in that a partial pre-alignment strategy dominated by an activation value is adopted, the maximum index of the activation value is dynamically counted, the mantissa of the maximum index is aligned, and multiple partial pre-alignment intermediate results are pre-calculated and latched for reuse; in combination with a customized lookup table and a multiplexer, a pre-calculation result is directly selected to replace real-time multiplication and displacement; and through the reconfigurable hardware, the FP8MACFP8 high-precision operation is realized by utilizing the FP8MACFP4 unit combination. According to the method, complete online floating point multiplication and addition operation is realized, and excellent energy efficiency ratio and operation speed are obtained while high precision is kept.
Owner:FUDAN UNIVERSITY

Systems and methods for energy-efficient, bit-parallel, multiply-accumulate for artificial intelligence and deep neural networks

A system and method for providing a tunable floating-point multiply-accumulate (MAC) unit are disclosed. The unit maintains full arithmetic precision while enabling dynamic elimination of ineffectual computation through operand decomposition and selective activation of partial product generation logic. The disclosed MAC unit is suitable for drop-in replacement in existing deep-learning accelerators and improves energy efficiency without requiring architectural changes.
Owner:KAXIRAS STEFANOS +3

Floating-point number multiplier, calculation chip and floating-point number multiplication operation method

The invention relates to a floating-point number multiplier, a computing chip and a floating-point number multiplication operation method, and belongs to the technical field of computers. According to the floating-point number multiplier, floating-point multiplication of a target floating-point number type is performed on input data by utilizing multiplication mode indication, the floating-point number type is specified as the floating-point number type of floating-point numbers in the input data, the floating-point numbers in the input data are divided into data groups by preprocessing the input data, and the floating-point numbers in the input data are divided into the data groups based on a target multiplication mode. According to the invention, the floating point data in the same data set is subjected to floating point number type floating point number multiplication operation, so that the floating point number multiplier can perform floating point number multiplication operation of various floating point number types by specifying different multiplication modes for input data of different floating point number types; therefore, the operation requirements of floating-point number multiplication of various floating-point number types can be met.
Owner:北京凌川科技有限公司 +1

Hybrid precision MAC tree structure for maximizing memory bandwidth usage to accelerate operation of generative large-scale language models

The present invention provides a hybrid precision MAC (multiple-and-multiple-accumulation) tree structure for maximizing memory bandwidth usage to accelerate the operation of a generative large scale language model. The present invention provides a hybrid precision MAC (multiple-and-multiple-accumulation) tree structure for maximizing memory bandwidth usage to accelerate the operation of the generative large scale language model. A MAC tree-based arithmetic unit according to one embodiment may include: a plurality of floating point multipliers connected in parallel and processing multiplication of data transferred from an external memory; a plurality of first converters for converting the output of each of the plurality of floating point multipliers from a floating point to a fixed point; a fixed-point adder tree which is connected to the plurality of first converters and processes the addition of the multiplication results of the plurality of floating-point multipliers; a fixed-point accumulator that accumulates the output of the fixed-point adder tree; and a second converter that converts the output of the fixed-point accumulator from a fixed point to a floating point.
Owner:超速有限公司

Floating point multiplications

Systems, apparatuses, and methods are disclosed for improved matrix-vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values, and a mode decoding unit is configured to provide a mode of the VMM operation according to a first floating point format of the activation values and a second floating point format of the weight values. A column cell of the CIM macro can be configured to output a product between an activation value and a weight value using a half adder and multiple multiplexers that provide selections to a full adder based on control signals.
Owner:OPENAI OPCO LLC

Floating point multiplication and addition instruction fusion method and device, equipment and storage medium

The invention relates to the technical field of computers, and provides a floating point multiplication and addition instruction fusion method and device, equipment and a storage medium. According to the method, a fusion scaling factor and a fusion offset of two continuous floating point multiplication and addition operations are calculated, the two continuous floating point multiplication and addition operations are merged based on the fusion scaling factor and the fusion offset, and a fusion instruction is output. And two FMA operations are combined into a single instruction by utilizing mathematical equivalent transformation, so that the problems of relatively large computing load, relatively high computing delay and low hardware utilization rate of a GPU (Graphic Processing Unit) caused by relatively large calling of floating point multiply-add instructions in exponential approximate calculation in existing softmax are solved.
Owner:GUANGZHOU WERIDE TECH LTD CO

Hardware circuit, method, device and medium for floating point matrix multiplication

The invention is suitable for the field of chip and digital circuit design, and provides a hardware circuit, a method, equipment and a medium for floating point matrix multiplication, and the hardware circuit comprises a floating point multiplication unit which is used for carrying out multiplication on an obtained multiplier in a first standard floating point format and an obtained multiplicand in a second standard floating point format, outputting a product expressed in a first user-defined floating point format; the floating point accumulation unit is used for adding the product and an accumulation input value and outputting an accumulation sum expressed by adopting a second user-defined floating point format; and the format conversion unit is used for converting the accumulated sum of the second custom floating point format into a third standard floating point format and outputting the third standard floating point format. According to the scheme, balance can be achieved between keeping of calculation precision and limiting of hardware overhead.
Owner:GUANGDONG LEAPFIVE TECH CO LTD

Method and apparatus for implied bit handling in floating point multiplication

Devices and methods are provided for performing, by a processor in response to a floating point multiply instruction, multiplication of floating point numbers. An example processor includes first, second, third, and fourth computational paths. In operation, the first determines values of implied bits of mantissas of floating point numbers and generates first partial product terms, the second multiplies remainders of the mantissas to generate second partial product terms, the third detects a number of leading zeros in the mantissas and determines a shift amount for each of the mantissas, and the fourth calculates exponents for a flush-to-zero mode.
Owner:TEXAS INSTRUMENTS INC

Floating point multiplications

Systems, apparatuses, and methods are disclosed for improved matrix–vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values, and a mode decoding unit is configured to provide a mode of the VMM operation according to a first floating point format of the activation values and a second floating point format of the weight values. A column cell of the CIM macro can be configured to output a product between an activation value and a weight value using a half adder and multiple multiplexers that provide selections to a full adder based on control signals.
Owner:OPENAI OPCO LLC

Tensor arithmetic unit based on RISC-V instruction set and intelligent processor

The invention provides a tensor arithmetic unit based on an RISC-V instruction set, and the tensor arithmetic unit comprises a microinstruction splitting and scheduling unit which receives a decoded macroscopic tensor instruction and an operation size parameter thereof, splits the macroscopic tensor instruction into microinstruction sequences according to a fixed physical scale of a reconfigurable calculation array, and transmits the microinstruction sequences to the reconfigurable calculation array; a zigzag traversal sequence and microinstruction scheduling are achieved through quintuple hardware circulation, and the loading, using and replacing sequence of the data blocks in the block register array is planned; the reconfigurable computing array is used for executing matrix multiply-accumulate operation of various data formats by designing a reconfigurable data path and fusing and multiplexing a floating point multiplier and a multi-precision accumulation tree under data paths with different bit widths; and the hardware multi-buffer unit is a multi-buffer architecture automatically managed by hardware and is matched with a zigzag traversal sequence to realize pipeline overlapping of data prefetching and calculation execution. The invention further provides an intelligent processor. Therefore, the invention can efficiently execute tensor operation with variable precision and variable scale.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Mixed-precision floating-point operation circuit in a dedicated processing block

Mixed precision floating point operation circuits in special purpose processing blocks are disclosed. Embodiments relate to integrated circuits having circuits that efficiently perform mixed precision floating point arithmetic operations. Such circuits can be implemented in special purpose processing blocks. The special purpose processing blocks can include configurable interconnect circuits to support a variety of different usage modes. For example, the special purpose processing blocks can implement fixed point addition, floating point addition, fixed point multiplication, floating point multiplication, sum of two multiplications, with or without conversion to a second floating point precision, followed by a subsequent addition in the second floating point precision if needed, to name a few. In some embodiments, two or more special purpose processing blocks can be arranged in a cascade chain and together perform more complex operations such as a recursive mode dot product of two vectors of floating point numbers having a first floating point precision and output the dot product in a second floating point precision.
Owner:ALTERA CORP

Floating point multiplication method and floating point multiplication circuit

The application provides a floating-point multiplication method and a floating-point multiplication circuit, and relates to the technical field of processors. The method comprises the following steps: obtaining a shift result of the product of a first floating-point number and a second floating-point number; the shift result is the result of a shift operation on the decimal part of the product; performing a mantissa rounding operation on the shift result to obtain a mantissa rounding result of the shift result; obtaining a first judgment result of whether the mantissa satisfies a first preset rounding condition in the process of the mantissa rounding operation; if the first judgment result is that the mantissa satisfies the first preset rounding condition, obtaining a target mantissa of the multiplication result of the first floating-point number and the second floating-point number according to the mantissa rounding result. The floating-point multiplication method has the advantage of short operation period.
Owner:BEIJING INSTITUTE OF OPEN SOURCE CHIP

Distributed double-precision floating-point multiplication

The present embodiments relate to circuitry that efficiently performs double-precision floating-point multiplication operations, single-precision floating-point multiplication operations, and fixed-point multiplication operations. Such circuitry may be implemented in specialized processing blocks. If desired, each specialized processing block efficiently may perform a single-precision floating-point multiplication operation, and multiple specialized processing blocks may be coupled together to perform a double-precision floating-point multiplication operation. Inter-block signaling circuits may generate rounding information and propagate the rounding information together with partial product results from a current specialized processing block to another specialized processing block.
Owner:ALTERA CORP

Transcendental function computation system and method based on interpolation approximation, and chip and terminal device

Provided in the present invention are a transcendental function computation system and method based on interpolation approximation, and a chip and a terminal device. The method comprises: an input module inputting floating-point numbers having a preset number of bits; a single-precision floating-point number computation module performing transcendental function computation on the input floating-point numbers in the computation mode of single-precision floating-point multiplication, and outputting computation results in a half-precision floating-point format; a half-precision floating-point number computation module performing transcendental function computation on the input floating-point numbers in the computation mode of half-precision floating-point multiplication, and outputting computation results in a half-precision floating-point format; and an output module outputting the computation results. By means of a single-precision floating-point number computation module and a half-precision floating-point number computation module, high-precision and high-performance operations can be performed on single-precision floating-point numbers and half-precision floating-point numbers, such that the requirements of processor chips and their interface APIs are met, thereby enabling the transcendental function computation method to not only reduce costs while ensuring precision, but also be applicable to various types of transcendental functions.
Owner:VERISILICON MICROELECTRONICS (SHANGHAI) CO LTD +4

Method and apparatus for implied bit handling in floating point multiplication

Devices and methods are provided for performing, by a processor in response to a floating point multiply instruction, multiplication of floating point numbers. In an example, a device includes a processor that includes a multiply circuit. The multiply circuit is configured to multiply floating point numbers in response to a floating point multiply instruction, and is further configured to determine values of implied bits of mantissas of the floating point numbers, and multiply the mantissas in parallel with the determining operation.
Owner:TEXAS INSTRUMENTS INC

Novel hybrid precision convolution multiplier-adder

The invention discloses a novel hybrid precision convolution multiplier-adder, and belongs to the technical field of artificial intelligence chip design and deep learning acceleration. The multiplier and adder supports two core data formats of INT16 and FP16, and efficient sharing of integer and floating point multiplication and addition resources is realized by improving a floating point number operation method. The core of the method is to deform a traditional FP16 floating-point number representation method, so that a 16-bit * 16-bit integer multiplier can be used in the calculation process of the method, and meanwhile, a multi-stage calculation architecture (including modules such as MTS calculation, product and maximum order code solution, symbol processing and shifting, adder tree summation and the like) is adopted, so that delay caused by multiple alignment of order codes is reduced. According to the design, multiplier and adder tree resources are multiplexed, on the premise that the accuracy loss is controllable, hardware resource redundancy and operation complexity are reduced, the advantage of comprehensive performance is more remarkable along with increase of the order number of the multiplier and adder, and the method is suitable for efficient acceleration scenes of deep learning convolution operation.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Hardware circuit, method, device and medium for floating point matrix multiplication operation

The application is suitable for the field of chip and digital circuit design, and provides a hardware circuit, method, equipment and medium for floating point matrix multiplication, wherein the hardware circuit comprises: a floating point multiplication unit, which is used for multiplying a first standard floating point format multiplier and a second standard floating point format multiplicand, and outputs a product in a first self-defined floating point format; a floating point accumulation unit, which is used for adding the product and an accumulation input value, and outputs an accumulation sum in a second self-defined floating point format; and a format conversion unit, which is used for converting the accumulation sum in the second self-defined floating point format into a third standard floating point format and then outputting. The scheme can balance between maintaining calculation accuracy and limiting hardware overhead.
Owner:GUANGDONG LEAPFIVE TECH CO LTD