Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

9 results about "Floating point multiplication" patented technology

Low-bit-width high-energy-efficiency floating point storage and calculation integrated circuit based on partial pre-alignment architecture

The invention belongs to the technical field of storage and calculation integration, and particularly relates to a low-bit-width and high-energy-efficiency floating point storage and calculation integrated circuit based on a partial pre-alignment framework. The circuit comprises a memory array, a pre-calculation unit, an adder tree, a configurable arithmetic unit and a normalization unit, and supports mixed precision operation of FP8MACFP4 and FP8MACFP8. The method is characterized in that a partial pre-alignment strategy dominated by an activation value is adopted, the maximum index of the activation value is dynamically counted, the mantissa of the maximum index is aligned, and multiple partial pre-alignment intermediate results are pre-calculated and latched for reuse; in combination with a customized lookup table and a multiplexer, a pre-calculation result is directly selected to replace real-time multiplication and displacement; and through the reconfigurable hardware, the FP8MACFP8 high-precision operation is realized by utilizing the FP8MACFP4 unit combination. According to the method, complete online floating point multiplication and addition operation is realized, and excellent energy efficiency ratio and operation speed are obtained while high precision is kept.
Owner:FUDAN UNIVERSITY

Hybrid precision MAC tree structure for maximizing memory bandwidth usage to accelerate operation of generative large-scale language models

The present invention provides a hybrid precision MAC (multiple-and-multiple-accumulation) tree structure for maximizing memory bandwidth usage to accelerate the operation of a generative large scale language model. The present invention provides a hybrid precision MAC (multiple-and-multiple-accumulation) tree structure for maximizing memory bandwidth usage to accelerate the operation of the generative large scale language model. A MAC tree-based arithmetic unit according to one embodiment may include: a plurality of floating point multipliers connected in parallel and processing multiplication of data transferred from an external memory; a plurality of first converters for converting the output of each of the plurality of floating point multipliers from a floating point to a fixed point; a fixed-point adder tree which is connected to the plurality of first converters and processes the addition of the multiplication results of the plurality of floating-point multipliers; a fixed-point accumulator that accumulates the output of the fixed-point adder tree; and a second converter that converts the output of the fixed-point accumulator from a fixed point to a floating point.
Owner:超速有限公司

Floating point multiplications

Systems, apparatuses, and methods are disclosed for improved matrix-vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values, and a mode decoding unit is configured to provide a mode of the VMM operation according to a first floating point format of the activation values and a second floating point format of the weight values. A column cell of the CIM macro can be configured to output a product between an activation value and a weight value using a half adder and multiple multiplexers that provide selections to a full adder based on control signals.
Owner:OPENAI OPCO LLC

Method and apparatus for implied bit handling in floating point multiplication

Devices and methods are provided for performing, by a processor in response to a floating point multiply instruction, multiplication of floating point numbers. An example processor includes first, second, third, and fourth computational paths. In operation, the first determines values of implied bits of mantissas of floating point numbers and generates first partial product terms, the second multiplies remainders of the mantissas to generate second partial product terms, the third detects a number of leading zeros in the mantissas and determines a shift amount for each of the mantissas, and the fourth calculates exponents for a flush-to-zero mode.
Owner:TEXAS INSTRUMENTS INC

Floating point multiplications

Systems, apparatuses, and methods are disclosed for improved matrix–vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values, and a mode decoding unit is configured to provide a mode of the VMM operation according to a first floating point format of the activation values and a second floating point format of the weight values. A column cell of the CIM macro can be configured to output a product between an activation value and a weight value using a half adder and multiple multiplexers that provide selections to a full adder based on control signals.
Owner:OPENAI OPCO LLC

Tensor arithmetic unit based on RISC-V instruction set and intelligent processor

The invention provides a tensor arithmetic unit based on an RISC-V instruction set, and the tensor arithmetic unit comprises a microinstruction splitting and scheduling unit which receives a decoded macroscopic tensor instruction and an operation size parameter thereof, splits the macroscopic tensor instruction into microinstruction sequences according to a fixed physical scale of a reconfigurable calculation array, and transmits the microinstruction sequences to the reconfigurable calculation array; a zigzag traversal sequence and microinstruction scheduling are achieved through quintuple hardware circulation, and the loading, using and replacing sequence of the data blocks in the block register array is planned; the reconfigurable computing array is used for executing matrix multiply-accumulate operation of various data formats by designing a reconfigurable data path and fusing and multiplexing a floating point multiplier and a multi-precision accumulation tree under data paths with different bit widths; and the hardware multi-buffer unit is a multi-buffer architecture automatically managed by hardware and is matched with a zigzag traversal sequence to realize pipeline overlapping of data prefetching and calculation execution. The invention further provides an intelligent processor. Therefore, the invention can efficiently execute tensor operation with variable precision and variable scale.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Transcendental function computation system and method based on interpolation approximation, and chip and terminal device

Provided in the present invention are a transcendental function computation system and method based on interpolation approximation, and a chip and a terminal device. The method comprises: an input module inputting floating-point numbers having a preset number of bits; a single-precision floating-point number computation module performing transcendental function computation on the input floating-point numbers in the computation mode of single-precision floating-point multiplication, and outputting computation results in a half-precision floating-point format; a half-precision floating-point number computation module performing transcendental function computation on the input floating-point numbers in the computation mode of half-precision floating-point multiplication, and outputting computation results in a half-precision floating-point format; and an output module outputting the computation results. By means of a single-precision floating-point number computation module and a half-precision floating-point number computation module, high-precision and high-performance operations can be performed on single-precision floating-point numbers and half-precision floating-point numbers, such that the requirements of processor chips and their interface APIs are met, thereby enabling the transcendental function computation method to not only reduce costs while ensuring precision, but also be applicable to various types of transcendental functions.
Owner:VERISILICON MICROELECTRONICS (SHANGHAI) CO LTD +4

Novel hybrid precision convolution multiplier-adder

The invention discloses a novel hybrid precision convolution multiplier-adder, and belongs to the technical field of artificial intelligence chip design and deep learning acceleration. The multiplier and adder supports two core data formats of INT16 and FP16, and efficient sharing of integer and floating point multiplication and addition resources is realized by improving a floating point number operation method. The core of the method is to deform a traditional FP16 floating-point number representation method, so that a 16-bit * 16-bit integer multiplier can be used in the calculation process of the method, and meanwhile, a multi-stage calculation architecture (including modules such as MTS calculation, product and maximum order code solution, symbol processing and shifting, adder tree summation and the like) is adopted, so that delay caused by multiple alignment of order codes is reduced. According to the design, multiplier and adder tree resources are multiplexed, on the premise that the accuracy loss is controllable, hardware resource redundancy and operation complexity are reduced, the advantage of comprehensive performance is more remarkable along with increase of the order number of the multiplier and adder, and the method is suitable for efficient acceleration scenes of deep learning convolution operation.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Hardware circuit, method, device and medium for floating point matrix multiplication operation

The application is suitable for the field of chip and digital circuit design, and provides a hardware circuit, method, equipment and medium for floating point matrix multiplication, wherein the hardware circuit comprises: a floating point multiplication unit, which is used for multiplying a first standard floating point format multiplier and a second standard floating point format multiplicand, and outputs a product in a first self-defined floating point format; a floating point accumulation unit, which is used for adding the product and an accumulation input value, and outputs an accumulation sum in a second self-defined floating point format; and a format conversion unit, which is used for converting the accumulation sum in the second self-defined floating point format into a third standard floating point format and then outputting. The scheme can balance between maintaining calculation accuracy and limiting hardware overhead.
Owner:GUANGDONG LEAPFIVE TECH CO LTD