Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

30 results about "Block floating-point" patented technology

Block floating point (BFP) is a method used to provide an arithmetic approaching floating point while using a fixed-point processor. The algorithm will assign an entire block of data an exponent, rather than single units themselves being assigned an exponent, thus making them a block, rather than a simple floating point. Block floating-point algorithm operations are done through a block using a common exponent, and can be advantageous to limit the space use in the hardware to perform the same functions as floating-point algorithms.

Apparatus using in-memory compute chiplet devices for inference-time compute acceleration

An apparatus using in-memory compute (IMC) chiplet devices for inference-time compute acceleration. The apparatus is configured to accelerate the workload computations for neural network models, such as those for Large Language Models (LLMs) and reasoning models. The apparatus achieves high throughput and low latency using a chiplet design, digital IMC (DIMC) based engines, efficient die-to-die (D2D) interconnects, block floating point (BFP) numerics, and large high bandwidth on-chip memories. With modular chiplets and efficient interconnects, the accelerator apparatus can be easily scaled to accelerate workloads for models of different sizes. The DIMC configuration within the chiplet slices also improves computational performance and reduces power consumption by integrating computational functions and memory fabric. And by dynamically switching between precision levels based on real-time analysis of a target workload, computational efficiency can be optimized while maintaining the necessary level of accuracy for each step of the workload computation.
Owner:D-MATRIX CORP

Acceleration method for accelerating multiplication and addition operation in large language model and hardware accelerator

The invention relates to the technical field of acceleration operation, in particular to an acceleration method for accelerating multiplication and addition operation in a large language model and a hardware accelerator, and the method comprises the following steps: dividing data in the form of m floating-point numbers into a floating-point number block, m being a preset positive integer; performing block floating point coding on the floating-point number block to obtain a shared index and a mantissa of the floating-point number block, and expanding bit widths of the shared index and the mantissa; performing multiplication and addition calculation based on the expanded sharing index and mantissa; the multiply-add calculation result is decoded and restored into a floating-point number form, and a decoding result is obtained; according to the acceleration method for accelerating the multiplication and addition operation in the large language model and the hardware accelerator, relatively high calculation precision can be kept.
Owner:GUANGDONG UNIV OF TECH

Bidirectional block floating point-based large language model reasoning acceleration method

The invention is applicable to the technical field of computers, and provides a big language model reasoning acceleration method based on bidirectional block floating points, which comprises the following steps: encoding an input text into a Token sequence, representing hidden representation and logits in a bidirectional block floating point number format in a Transform reasoning process, and combining a Softmax normalization method of a table look-up method based on bidirectional block floating points to obtain a big language model reasoning acceleration model; and low-bit efficient reasoning is realized, and a reasoning result is generated. According to the method, the calculation complexity and the storage overhead of the reasoning process are remarkably reduced while the generation precision is kept, and the reasoning speed and the energy efficiency ratio of the large language model are improved.
Owner:NANJING INST OF TECH

Stacked apparatus using in-memory compute chiplet devices for inference-time compute acceleration

A stacked apparatus using in-memory compute (IMC) chiplet devices for inference-time compute acceleration. The apparatus is configured to accelerate the workload computations for neural network models, such as those for Large Language Models (LLMs) and reasoning models. The apparatus achieves high throughput and low latency using a chiplet design, digital IMC (DIMC) based engines, efficient die-to-die (D2D) interconnects, block floating point (BFP) numerics, and large high bandwidth on-chip memories. With modular chiplets in stacked configurations with memory devices and efficient interconnects, the accelerator apparatus can be easily scaled to accelerate workloads for models of different sizes. The DIMC configuration within the chiplet slices also improves computational performance and reduces power consumption by integrating computational functions and memory fabric. And by dynamically switching between precision levels based on real-time analysis of a target workload, computational efficiency can be optimized while maintaining accuracy.
Owner:D-MATRIX CORP

Task processing method, device and equipment based on adaptive block floating point data

The invention relates to the technical field of task processing, and discloses a task processing method, device and equipment based on adaptive block floating point data, and the method comprises the steps: receiving a group of floating point data corresponding to a target task as a current data block; determining a block floating point parameter of the current data block according to a preset rule or at least one data characteristic of the current data block; calculating a sharing index for the current data block according to a sharing index determination strategy; processing each floating point data in the current data block based on the sharing index to obtain respective corresponding mantissas; combining the sharing index and each mantissa into a block floating point representation of the current data block; and processing the target task based on the block floating point representation of the current data block. According to the task processing method and device, the problem that the processing precision and the processing efficiency are unbalanced when task processing is carried out based on block floating point data representation in the prior art is solved, and the balance of the task processing precision and the task processing efficiency can be ensured.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Dynamic range channelization receiver method and system based on block floating point and AGC

The invention relates to the technical field of channelized receiver management, in particular to a dynamic range channelized receiver method and system based on a block floating point and AGC (Automatic Gain Control), in the system, an automatic gain control module is used for detecting the power of an obtained intermediate frequency analog signal, generating an AGC voltage corresponding to the intermediate frequency analog signal and feeding back the AGC voltage to an analog front end; and dynamically updating the gain of the variable gain amplifier in the radio frequency signal preprocessing process. According to the invention, each channel has independent block floating point gain control, and the channels do not interfere with each other, so that parallel signal-to-noise ratio processing is realized; meanwhile, a complete gain control history is constructed by jointly recording the AGC gain and block floating point processing, data support is provided for recovering the original absolute power value of the signal in each channel, and the contradiction between gain control and information reservation is solved to a certain extent.
Owner:NANJING NAT ELECTRONIC TECH CO LTD

Block floating point calculation device and differential equation calculation system

The invention provides a block floating point arithmetic device and a differential equation calculation system, and relates to the field of calculation devices, and the differential equation calculation system manages the flow of data organized in a differential format among storage hierarchies through a buffer controller. And a processing unit array synchronously executes Gaussian-Seidel iterative calculation for differential format optimization based on a wavefront sequence, then compresses a result through a block floating point quantizer, and outputs a solution matrix through a differential reduction unit after convergence, a plurality of parallel processing units in the processing unit array combine precision configuration with a block floating point input format to form a collaborative optimization calculation mode, the complexity of mantissa processing is further controlled in the calculation process by sharing index compression data and dynamic precision parameter permission, and the problems of resource occupation requirements, memory access times and high energy consumption are solved.
Owner:NANJING UNIV

Data coding method and device based on block chain, equipment and storage medium

The invention discloses a data coding method and device based on a block chain, equipment and a storage medium. The data coding method comprises the following steps: reading floating point data in a block body of a first block; encoding the floating point data to obtain block floating point data; and storing a sign bit and a mantissa bit of the block floating point data into a block body of the first block, and storing a sharing index of the block floating point data into a block head of the first block. According to the application, the floating point data in the block body of the first block is encoded to obtain the block floating point data; the sign bit and the mantissa bit of the block floating point data are stored in the block body of the first block, and the sharing index of the block floating point data is stored in the block head of the first block, so that the storage space of the index bit is released, the on-chain storage overhead and the network transmission load of the block chain can be reduced, the data transmission speed is improved, and the user experience is improved. And the performance of the block chain is optimized.
Owner:SHENZHEN LINZHOU TECHNOLOGY CO LTD

A training method, device and computing device for a neural network model

ActiveCN113570053BNeural learning methodsForward propagationBlock floating-point
The present invention discloses a training method, apparatus, and computing device for a neural network model. The method includes a forward propagation step, a backpropagation step, and a parameter update step. In the parameter update step, the following steps are performed: calculating parameter update values for the current network layer based on parameter gradients, generating a fourth block of floating-point numbers corresponding to the parameter update values, wherein the bit width of the fourth block of floating-point numbers is a third predetermined value; and updating the parameters of the current network layer based on the first block of floating-point numbers and the fourth block of floating-point numbers, generating a first block of floating-point numbers corresponding to the updated parameters.
Owner:T-HEAD (SHANGHAI) SEMICON CO LTD

Neural network activation compression with non-uniform mantissa

Apparatuses and methods for training a neural network accelerator using quantization precision data formats are disclosed, and in particular for storing activation values from a neural network in a compressed format having lossy or non-uniform mantissa for use during forward and backward propagation training of the neural network. In certain examples of the disclosed technology, a computing system includes a processor, a memory, and a compressor in communication with the memory. The computing system is configured to perform forward propagation for a layer of a neural network to produce first activation values in a first block floating point format. In some examples, the activation values generated by the forward propagation are converted by the compressor to a second block floating point format having a non-uniform and / or lossy mantissa. The compressed activation values are stored in the memory, where they can be retrieved for use during backward propagation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Training neural network accelerators using mixed precision data formats

Techniques related to training neural network accelerators using mixed precision data formats are disclosed. In one example of the disclosed technology, a neural network accelerator is configured to accelerate a given layer of a multi-layer neural network. The input tensor of a given layer can be converted from a common precision floating point format to a quantized precision floating point format. A tensor operation may be performed using the converted input tensor. The tensor operation result can be converted from a block floating point format to a common precision floating point format. The converted result can be used to generate an output tensor of the layer of the neural network, wherein the output tensor is in a normal precision floating point format.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Enhanced block floating point compression for open radio access network fronthaul

A distributed unit (DU) signals a maximum IQ data bit width for downlink communications associated with a zone identifier (ID) to a radio unit (RU). The DU signals a per physical resource block (PRB) bit width parameter for downlink communications to the RU. The DU transmits the downlink communications based on at least one of the maximum IQ data bit width or the bit width parameter for the PRB, and the RU receives the downlink communications based on at least one of the maximum IQ data bit width or the bit width parameter for the PRB. For uplink communications, the DU transmits a first indication of a maximum IQ data bit width in a control plane message to the RU. The RU transmits a second indication of a per PRB bit width parameter for uplink communications to the DU. The RU transmits the uplink communications based on at least one of the maximum IQ data bit width or the bit width parameter for the PRB, and the DU receives the uplink communications based on at least one of the maximum IQ data bit width or the bit width parameter for the PRB.
Owner:QUALCOMM INC

Cross-chain data sending method and receiving method, electronic equipment and storage medium

The invention discloses a cross-chain data sending method and receiving method, electronic equipment and a storage medium. The cross-chain data sending method comprises the following steps: performing block coding on first floating point data to be transmitted to obtain first block floating point data; performing structured packaging on the first block of floating point data to obtain a second block of floating point data; and performing cross-chain transmission on the second block of floating point data to send the second block of floating point data to the second block chain. In a cross-chain scene, data is compressed through block coding, the on-chain storage overhead and the network transmission load of the block chain are reduced, and the data transmission speed and the cross-chain communication efficiency are improved; through structured packaging, the problem of inconsistent analysis of traditional floating point data in a multi-chain environment is avoided, and the exchange efficiency and consistency of the floating point data between multi-chain systems are improved. The method is suitable for cross-chain floating point calculation or data synchronization scenes, in particular to traceability chains, cross-chain model calculation, credible floating point data verification and other scenes with low data precision requirements but large data volume.
Owner:SHENZHEN LINZHOU TECHNOLOGY CO LTD

Floating point tensor data auditing method and device, equipment and storage medium

The invention discloses a floating point tensor data auditing method and device, equipment and a storage medium. The floating point tensor data auditing method comprises the following steps: dividing floating point tensor data to be processed into a plurality of blocks with fixed sizes; performing block coding on the floating point tensor data in each block to obtain block floating point tensor data; and in a block chain node, auditing the block floating point tensor data. According to the method, the floating-point tensor data is subjected to block processing, the floating-point tensor data is compressed and calculated in a block floating-point representation mode, and the data scale is remarkably compressed, so that the on-chain storage and bandwidth overhead is reduced, and the on-chain auditing efficiency is improved. The method can be implemented on the intelligent contract and consensus mechanism of the existing block chain platform (such as Ethereum, Fabric and the like), and has relatively high landing capability.
Owner:SHENZHEN LINZHOU TECHNOLOGY CO LTD

A hardware implementation method and device of a GAN network, a storage medium and a terminal

The application discloses a kind of hardware implementation method, device, storage medium and terminal of GAN network, method includes: training GAN network obtains the floating-point number model after convergence, and exports the floating-point number model after convergence;According to the network parameter of GAN network, floating-point number model is converted into block floating-point model;The block floating-point convolution structure of block floating-point model is deployed on hardware, and convolution operation is carried out based on the block floating-point convolution structure after deployment.The hardware implementation method provided in the application has the advantages of simple implementation steps, less network parameter precision loss, low operation complexity and significantly reduced storage unit requirements, solves the problem that existing generative adversarial network has large resource overhead and slow inference speed when deployed on hardware.
Owner:INST OF MICROELECTRONICS CHINESE ACAD OF SCI LTD

Acceleration chip, data processing method and application

The invention discloses an acceleration chip which comprises an off-chip storage unit used for storing input activation data and weight matrix data of a block floating point data structure; the on-chip cache unit is used for caching block floating point data to be calculated; the block floating point number direct carrying unit is used for carrying block floating point data between the off-chip storage unit and the on-chip cache unit; the block floating-point number matrix multiply-add unit is used for executing mantissa fixed-point multiply-add operation based on the block floating-point data structure and generating a floating-point result; a data format conversion unit for converting between a floating point format and a block floating point data structure; and the scheduling control unit is used for coordinating the work of the block floating point number direct carrying unit, the block floating point matrix multiplication and addition unit and the data format conversion unit. The invention further discloses a data processing method which has a wide application prospect.
Owner:SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD

Neural network device performing floating point operation and operating method thereof

ActiveCN113807493BDigital data processing detailsAnalogue-digital convertersComputer hardwareMultiply–accumulate operation
A neural network device performing floating point operations and an operating method thereof are provided. The neural network device performs a multiply-accumulate (MAC) operation for a product of a fraction of a weight and an input activation in a block floating point format by using an analog crossbar array, performs an addition operation for a shared exponent of the weight and the input activation in the block floating point format by using a digital computing circuit, and outputs a partial sum of a floating point output activation by combining a result of the MAC operation with a result of the addition operation.
Owner:SAMSUNG ELECTRONICS CO LTD

System and method for micromachine learning using block floating points

A system includes a first FP to BFP converter, a second FP to BFP converter, an 8-bit integer multiplier, an adder, an accumulator, and a BFP to FP converter. The first FP-to-BFP converter and the second FP-to-BFP converter respectively receive 32-bit floating point pixels and filter data and simplify mantissas of the 32-bit floating point pixels and the filter data into an 8-bit BFP format. The 8-bit integer multiplier processes the values in the BFP format through multiply-accumulate operation and generates data of a 16-bit product. An adder accumulates a plurality of 16-bit multiplied data into 64-bit sum data, and an accumulator further aggregates them. The BFP-to-FP converter converts the data of the 64-bit cumulative sum into output data represented by 32-bit FP.
Owner:HONG KONG APPLIED SCI & TECH RES INST

BLOCK FLOATTING POINT CALCULATIONS WITH COMMON EXPONS

ActiveDE602019086343T2Block floating-pointStructural engineering
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Neural network activation compression with non-uniform mantissas

Apparatuses and methods are disclosed for training neural network accelerators using quantized precision data formats, and in particular for storing activation values from neural networks in a compressed format with lossy or non-uniform mantissas for use during forward and backward propagation training of neural networks. In some examples of the disclosed technology, a computing system includes a processor, a memory, and a compressor in communication with the memory. The computing system is configured to perform forward propagation for layers of a neural network to generate a first activation value in a first block floating point format. In some examples, an activation value generated by forward propagation is converted by a compressor into a second block floating point format having a non-uniform and / or lossy mantissa. The compressed activation value is stored in a memory in which the compressed activation value may be retrieved for use during backward propagation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Neural network activation compression with narrow block floating-point

Apparatus and methods for training a neural network accelerator using quantized precision data formats are disclosed, and in particular for storing activation values from a neural network in a compressed format for use during forward and backward propagation training of the neural network. In certain examples of the disclosed technology, a computing system includes processors, memory, and a compressor in communication with the memory. The computing system is configured to perform forward propagation for a layer of a neural network to produced first activation values in a first block floating-point format. In some examples, activation values generated by forward propagation are converted by the compressor to a second block floating-point format having a narrower numerical precision than the first block floating-point format. The compressed activation values are stored in the memory, where they can be retrieved for use during back propagation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Block floating point operation device and differential equation calculation system

The application provides a block floating point operation device and a differential equation calculation system, and relates to the field of calculation devices. The differential equation calculation system manages the flow of data organized in a difference format between storage levels through a buffer controller, synchronously performs Gauss-Seidel iteration calculation optimized for the difference format based on a wavefront sequence by a processing unit array, compresses the result through a block floating point quantizer, and outputs a solution matrix through a difference reduction unit after convergence. In the processing unit array, multiple parallel processing units combine precision configuration with a block floating point input format to form a cooperatively optimized calculation mode. By sharing index compression data, dynamic precision parameters allow further control of the complexity of the mantissa during the calculation process, solving the problems of high resource occupation demand, memory access frequency and energy consumption.
Owner:NANJING UNIV

Data compression method and device, data decompression method and device, communication device and medium

The invention provides a data compression method and device, a data decompression method and device, a communication device and a medium, which can be applied to the technical field of communication. The method is applied to a first communication device and comprises the following steps: performing block floating point compression on first data to be transmitted to obtain a plurality of index factors and mantissa data corresponding to the index factors respectively; performing decimal bit quantization on the mantissa data to obtain a plurality of quantized data blocks; the bit width of the quantized data included in the quantized data block is smaller than that of the mantissa data; second data is transmitted to a second communication device, the second data including the plurality of exponential factors and the plurality of quantized data blocks. The compression mode provided by the invention has good compression capability, reduces the transmission bandwidth demand, and improves the bandwidth utilization rate.
Owner:RUIJIE NETWORKS CO LTD

Compression and storage of neural network activations for backpropagation

Apparatus and methods for training a neural network accelerator using quantized precision data formats are disclosed, and in particular for storing activation values from a neural network in a compressed format for use during forward and backward propagation training of the neural network. In certain examples of the disclosed technology, a computing system includes processors, memory, and a compressor in communication with the memory. The computing system is configured to perform forward propagation for a layer of a neural network to produced first activation values in a first block floating-point format. In some examples, activation values generated by forward propagation are converted by the compressor to a second block floating-point format having a narrower numerical precision than the first block floating-point format. The compressed activation values are stored in the memory, where they can be retrieved for use during back propagation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Floating point data processing method and system

PendingCN120872286ADigital data processing detailsNumerical stabilityBlock floating-point
The invention relates to the field of artificial intelligence calculation and deep learning acceleration, and discloses a floating point data processing method and system, and the method comprises the following steps: presetting a quantized index bit width le, and dividing the index space of floating point data into a normal value interval and an abnormal value interval; dividing every m floating-point numbers into one data block, and calculating and storing block floating-point data of each index fold including a scaling factor alpha and a fold index Efold based on a normal value interval and an abnormal value interval; and carrying out multiplication and addition operation on the block floating point data, and outputting updated floating point data. According to the method, the problem of insufficient numerical stability in the prior art is solved, and the method has the characteristics of ensuring the calculation precision and reducing the calculation resource consumption at the same time.
Owner:GUANGDONG UNIV OF TECH

Dynamically Mixed Precision Machine Learning Systems and Methods

PendingUS20250224920A1Digital data processing detailsBlock floating-pointFeature data
Systems, methods, and circuitry for dynamically mixed precision machine learning are provided. An integrated circuit may include conversion circuitry to convert input feature data to block floating point format and upper / lower splitter circuitry to split the input feature data in the block floating point format into an upper component in the block floating point format and a lower component in the block floating point format. A processing element may use only the upper component when operating in a lower-precision mode and use both the upper component and the lower component when operating in a higher-precision mode.
Owner:KERTESZ AUDREY +4

Adjusting precision and topology parameters for neural network training based on a performance metric

Apparatus and methods for training neural networks based on a performance metric, including adjusting numerical precision and topology as training progresses are disclosed. In some examples, block floating-point formats having relatively lower accuracy are used during early stages of training. Accuracy of the floating-point format can be increased as training progresses based on a determined performance metric. In some examples, values for the neural network are transformed to normal precision floating-point formats. The performance metric can be determined based on entropy of values for the neural network, accuracy of the neural network, or by other suitable techniques. Accelerator hardware can be used to implement certain implementations, including hardware having direct support for block floating-point formats.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

System and method for tiny machine learning using block floating point

A system includes a first FP-to-BFP converter, a second FP-to-BFP converter, an 8-bit integer multiplier, an adder, an accumulator, and a BFP-to-FP converter. The first and second FP-to-BFP converters receive 32-bit floating-point pixel and filter data, respectively, reducing their mantissas to 8-bit BFP format. The 8-bit integer multiplier processes these BFP values via multiply-accumulate operations, generating a 16-bit product. The adder accumulates multiple 16-bit products into a 64-bit sum, which the accumulator further aggregates. The BFP-to-FP converter transforms the 64-bit accumulated sum into a 32-bit floating-point output.
Owner:HONG KONG APPLIED SCI & TECH RES INST