Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

13 results about "Single-precision floating-point format" patented technology

Single-precision floating-point format is a computer number format, usually occupying 32 bits in computer memory; it represents a wide dynamic range of numeric values by using a floating radix point. A floating-point variable can represent a wider range of numbers than a fixed-point variable of the same bit width at the cost of precision. A signed 32-bit integer variable has a maximum value of 2³¹ − 1 = 2,147,483,647, whereas an IEEE 754 32-bit base-2 floating-point variable has a maximum value of (2 − 2⁻²³) × 2¹²⁷ ≈ 3.4028235 × 10³⁸. All integers with 7 or fewer decimal digits, and any 2 for a whole number −149 ≤ n ≤ 127, can be converted exactly into an IEEE 754 single-precision floating-point value.

Circuit for approximate floating point fusion dot product operation

The invention discloses a circuit for approximate floating point fusion dot product operation. The circuit comprises an extraction module used for extracting sign bits, mantissa bits and exponent bits from four single-precision floating point numbers; the symbol integration module is used for integrating symbol bits and correcting mantissa bits; the index comparison module is used for calculating two groups of dot product effective indexes according to the index bits and generating three control quantities; the first multiplication module is used for generating a first group of approximate partial products according to a correction mantissa digit corresponding to a first control quantity larger value; the second multiplication module is used for generating a second group of approximate partial products according to the correction mantissa digit corresponding to the smaller value of the first control quantity, and performing arithmetic displacement and adding symbol compensation bits according to the size of the second control quantity; the fusion compression module is used for performing fusion compression on the first group of approximate partial products, the second group of approximate partial products after arithmetic shift and the symbol compensation bits to obtain a sum sequence and a carry sequence; and the result output module is used for generating a dot product operation result in combination with the third control quantity.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Multi-level virtual-real coordinate mapping and drift correction method, system and platform

PendingCN121722365ASoftware designVisual/graphical programmingVirtual coordinate systemsControl engineering
The invention provides a multi-level virtual-real coordinate mapping and drift correction method, system and platform, a dynamic floating original point mechanism is established, the original point position of a virtual coordinate system is adjusted in real time, so that a scene around a user is always in a high-precision calculation range of a single-precision floating-point number, and the user experience is improved. The dynamic floating origin mechanism comprises a coordinate mapping engine, an origin decision module and a seamless switching manager. A multi-level real world-virtual-real space-virtual environment conversion link is directly opened, large-scale development of augmented reality XR contents of cities / parks / outdoor / indoor / temporary spaces, automatic driving, unmanned aerial vehicles, intelligent robot navigation and the like is promoted, and the problems of large-space rendering jitter, low development efficiency and the like in the prior art are solved. In order to achieve the purpose, the invention provides an underlying OS operating system, and provides a cross-multi-level real world coordinate system and virtual environment coordinate system mapping method, a drift correction system and a developer operating platform.
Owner:WUHAN HUACHUANG HIGHLIGHT DIGITAL TECHNOLOGY CO LTD

Asymmetric multiply-accumulators, multiply-accumulator methods and electronic devices

This application relates to a multiply-accumulate method, apparatus, processor, and computer program product. The method includes: when the logic operation unit performs single-precision floating-point multiply-accumulate operations, two half-precision multiply-accumulate units in each single-precision multiply-accumulate unit combine to perform multiply-accumulate operations on the single-precision floating-point number to be processed, obtaining the corresponding single-precision multiply-accumulate result, for a total of N multiply-accumulate results; when the logic operation unit performs half-precision floating-point multiply-accumulate operations, each half-precision multiply-accumulate unit performs multiply-accumulate operations on the half-precision floating-point number to be processed, obtaining the corresponding half-precision multiply-accumulate result, for a total of 2N multiply-accumulate results. This improves the utilization rate of the multiply-accumulate units.
Owner:GLENFLY TECH CO LTD

On-chip storage protocol device and method for high-precision matrix multiplication

The invention provides an on-chip storage protocol device for high-precision matrix multiplication. The on-chip storage protocol device comprises a scheduling unit; the index and fragmentation preparation unit is used for reading single-precision floating point data from the on-chip first-level cache, performing index domain information extraction and low bit width conversion, generating a reference index and a low bit width fragmentation operand, and outputting the low bit width fragmentation operand and the reference index to the scheduling unit; the processing array is used for receiving the low-bit wide fragmentation operands and the reference indexes, executing matrix multiplication and generating multiple paths of intermediate multiplication results; the matrix reduction unit is used for receiving the multiple paths of intermediate multiplication results and the reference index, distributing the multiple paths of intermediate multiplication results to the corresponding memory banks according to the reference index and carrying out accumulation updating; and the normalization unit is used for receiving accumulated and updated results and carrying out floating point format reduction processing to generate final single-precision floating point data. According to the invention, the throughput and storage efficiency of the on-chip protocol are improved, and the hardware area and power consumption are reduced.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Systems and methods for performing 16-bit floating-point matrix dot product instructions

In an embodiment, an apparatus comprises a processor having a plurality of cores to perform operations corresponding to an instruction. The instruction is to identify a first storage location of a first matrix having M rows by K columns of pairs of 16-bit floating-point data elements having a bfloat16 format, a second storage location of a second matrix having K rows by N columns of pairs of 16-bit floating-point data elements having the bfloat16 format, and a third storage location of a third matrix having M rows by N columns of 32-bit single precision floating-point data elements, wherein the M rows of the first matrix and the N columns of the second matrix are specified in a register. The operations include to, for each data element position at a row m of the M rows and at a column n of the N columns of the third matrix: generate a dot product from K pairs of the 16-bit floating-point data elements corresponding to a row m of the first matrix and corresponding ones of K pairs of the 16-bit floating-point data elements corresponding to a column n of the second matrix; accumulate the dot product with a 32-bit single precision floating-point data element corresponding to the data element position to generate a result 32-bit single precision floating-point data element; and store the result 32-bit single precision floating-point data element in the data element position of the third matrix.
Owner:INTEL CORP

Mixing precision tensor calculation unit and method for edge calculation

The invention relates to the technical field of artificial intelligence hardware acceleration, in particular to a mixed precision tensor calculation unit and method for edge calculation. The reconfigurable computing architecture based on the systolic array is designed for the contradiction between edge side hardware resource limitation and model reasoning high-performance requirements. The system comprises an AXI bus interface, a global control module, a data formatting module and a multi-Bank parallel cache structure. Mixed precision calculation is supported, and when core matrix multiplication and addition operation is processed, a half-precision floating-point number is adopted as an input operand, and a single-precision floating-point number is adopted as an accumulation and output operand, so that the data throughput is maximized on the premise of ensuring the calculation precision. Through the optimized two-dimensional mesh topology systolic array and data flow control strategy, on-chip storage access conflicts can be effectively reduced, transmission delay is reduced, the balance of high computing power, high energy efficiency and low resource occupation is achieved, and the method is suitable for complex load prediction model reasoning in the fields of intelligent power grids and the like.
Owner:GANSU ELECTRIC POWER TIANSHUI POWER SUPPLY

Reconfigurable fractional order computing system with efficient use of fpga resources

The application discloses a reconfigurable fractional order calculation system with high-efficiency FPGA resource utilization. After input data is normalized and converted into single-precision floating-point numbers by a data preprocessing module, a control module receives binomial coefficient theoretical calculation parameters and binomial coefficient segmented linear fitting parameters set by a user, controls a binomial coefficient fitting module to calculate binomial coefficients and perform segmented linear fitting, configures configuration parameters required by a fixed window length calculation module and a segmented linear function calculation module according to a fitting result, and starts the fixed window length calculation module and the segmented linear function calculation module to perform fractional order calculation after the configuration is completed, so that a fractional order calculation result of the input data is obtained. The application is based on a fixed window (FWL) with error compensation and a multi-segment linear function (PWL), realizes a real-time reconfigurable fractional order calculation system on an FPGA platform, improves the FPGA resource utilization efficiency, and guarantees the precision and efficiency of the fractional order calculation.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Instructions for dual destination type conversion, accumulation, and atomic memory operations

Disclosed embodiments relate to instructions for double destination type conversion, accumulation, and atomic memory operations. In one example, a system includes a memory; a processor including an extraction circuit to extract an instruction from a code storage, the instruction including an opcode, a first destination identifier, and a source identifier to specify a source vector register, the source vector register including a plurality of single precision floating point data elements; a decode circuit to decode the extracted instruction; and an execution circuit to execute the decoded instruction to convert elements of the source vector register to double precision floating point values, store a first half of the double precision floating point values to a first location identified by the first destination identifier, and store a second half of the double precision floating point values to a second location.
Owner:INTEL CORP

LFMCW radiation source L array DBF direction finding method

The application discloses a kind of L array DBF goniometry methods for LFMCW radiation source, the data received by L array is down-converted and extracted filtering in FPGA then according to two-dimensional linear array processing, utilize the advantage of parallel processing of FPGA, when DBF processing, two parallel routes of search and processing are divided, then the idea of two linear arrays and difference beam is FFT processed using the IP of FFT, after the result is converted into single-precision floating-point number, the data in FPGA is transmitted to ARM using the AXI IP of Zynq for subsequent processing and angle conversion.The application utilizes the low data amount DBF processing method, realizes the efficient transmission of data between FPGA and ARM by combining the AXI data line of Zynq series, greatly improves the overall performance of signal processing machine, ensures basic goniometry function while enhancing the real-time processing capability of system.
Owner:NO 8511 RES INST OF CASIC

Single-precision floating point matrix multiplication calculation unit, splitting mapping method and accelerator

PendingCN121979486AImprove reuse efficiencyreduce overheadDigital data processing detailsEnergy efficient computingAlgorithmInteger matrix
The invention provides a single-precision floating-point matrix multiplication calculation unit, and the unit comprises a preprocessing module which is used for reading single-precision floating-point data, and dividing the mantissa field part of each matrix into a plurality of groups of integer matrixes with low bit width; the linear evaluation arithmetic logic module is used for executing linear combination operation on the multiple groups of low-bit-width integer matrixes to generate multiple groups of point value matrixes; the low-bit-width integer calculation array is used for executing low-bit-width integer matrix multiplication operation on the multiple groups of point value matrixes to obtain multiple groups of point value product matrixes; the interpolation merging module is used for executing interpolation operation and weighted merging on the multiple groups of point value product matrixes to generate a mantissa field product result; and the format conversion module is used for combining the mantissa field product result with the exponential field part of each matrix to generate a single-precision floating-point number matrix multiplication result. According to the invention, the overall throughput and energy efficiency are improved, and the mapping calculation and system overhead are reduced.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

A method for using Cuda to accelerate large-scale BA parallel optimization

This invention discloses a method for accelerating large-scale BA parallel optimization using CUDA, comprising the following steps: S1, constructing the energy residual equation of the system; S2, solving the linear equation system; S2-1, applying J... l Perform QR decomposition; S2-2 Use QR decomposition to marginalize the linear equation system; S3, Solve to obtain Δx p Δx l Substituting the original system equations, BA optimization is achieved. The above technical solution improves the numerical stability of the computation and allows for solving large-scale bundle adjustment problems using single-precision floating-point numbers, resulting in good parallelism in solving the linear equation system Ax = b. This significantly improves the parallelism of solving the linear equations, effectively increasing the efficiency of solving the linear equation system. This effectively overcomes the limitations of existing hardware resources, enabling optimization in the large-scale 3D scene reconstruction process, and greatly improving optimization efficiency, particularly for global pose and observation point optimization. The optimization results can serve as the basis for reconstructing sparse 3D point clouds of even larger-scale 3D scenes.
Owner:HANGZHOU DIANZI UNIV +2

Transcendental function computation system and method based on interpolation approximation, and chip and terminal device

Provided in the present invention are a transcendental function computation system and method based on interpolation approximation, and a chip and a terminal device. The method comprises: an input module inputting floating-point numbers having a preset number of bits; a single-precision floating-point number computation module performing transcendental function computation on the input floating-point numbers in the computation mode of single-precision floating-point multiplication, and outputting computation results in a half-precision floating-point format; a half-precision floating-point number computation module performing transcendental function computation on the input floating-point numbers in the computation mode of half-precision floating-point multiplication, and outputting computation results in a half-precision floating-point format; and an output module outputting the computation results. By means of a single-precision floating-point number computation module and a half-precision floating-point number computation module, high-precision and high-performance operations can be performed on single-precision floating-point numbers and half-precision floating-point numbers, such that the requirements of processor chips and their interface APIs are met, thereby enabling the transcendental function computation method to not only reduce costs while ensuring precision, but also be applicable to various types of transcendental functions.
Owner:VERISILICON MICROELECTRONICS (SHANGHAI) CO LTD +4

A method for quickly solving reciprocal of positive numbers based on SIMD instruction implementation

This invention provides a method for quickly calculating the reciprocal of a positive number based on SIMD instructions, comprising: S1. Loading data, which is loaded in multiples of 32 bits, with a maximum of 512 bits of data per register; single-precision floating-point numbers are 32 bits, and a register can load 16 floating-point numbers, so one SIMD instruction can load or calculate 16 floating-point numbers simultaneously. Let Register1 = Ingenic_simd512_load((float)data); Ingenic_simd512_load is the SIMD instruction for loading data; (float)data is the 16 input 32-bit single-precision floating-point numbers; Register1 is register 1, and the input data is stored in Register1. S2. Use SIMD instructions to calculate the reciprocal of the square root from the data in Register1, and store the result in Register2; S3. Square the reciprocal of the square root from step S2 to obtain the reciprocal. Let Register3 = Ingenic_simd512_float_mul(Register2, Register2); Ingenic_simd512_float_mul is a SIMD instruction for floating-point multiplication. Multiply the data in Register2 with the data in Register2, that is, square the data in Register2, and store the result in Register3; S4. Save the calculation result from the register to memory.
Owner:INGENIC SEMICON CO LTD