Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

27 results about "Floating-point unit" patented technology

A floating-point unit (FPU, colloquially a math coprocessor) is a part of a computer system specially designed to carry out operations on floating point numbers. Typical operations are addition, subtraction, multiplication, division, square root, and bitshifting. Some systems (particularly older, microcode-based architectures) can also perform various transcendental functions such as exponential or trigonometric calculations, though in most modern processors these are done with software library routines.

STORING FLOATING-POINT VALUES ACCORDING TO AN EXTENDED QFLOAT FLOATING-POINT (xqFP) FORMAT IN PROCESSOR DEVICES

Storing floating-point values according to an extended QFloat floating-point (xqFP) format in processor devices is disclosed herein. In some aspects, a processor device comprises a register file comprising a plurality of registers, and comprises a floating-point unit (FPU) circuit that is configured to store a first floating-point value in a register of the plurality of registers. The first floating-point value is formatted according to the xqFP format that comprises an exponent field and a significand field. The significand field is formatted as a signed one's complement value, and comprises a sign bit, an explicit most-significant-bit (MSB), a fractional field, and a deferred increment bit that represents a value of one-half (½) unit of least precision (ULP).
Owner:QUALCOMM INC

Register file for systolic array

A processing apparatus includes a general-purpose parallel processing engine including a set of multiple processing elements including a single precision floating-point unit, a double precision floating point unit, and an integer unit; a matrix accelerator including one or more systolic arrays; a first register file coupled with a first read control circuit, wherein the first read control circuit couples with the set of multiple processing elements and the matrix accelerator to arbitrate read requests to the first register file from the set of multiple processing elements and the matrix accelerator; and a second register file coupled with a second read control circuit, wherein the second read control circuit couples with the matrix accelerator to arbitrate read requests to the second register file from the matrix accelerator and limit access to the second register file by the set of multiple processing elements.
Owner:INTEL CORP

Performing floating-point operations using an expanded-range floating-point format in processor devices

Performing floating-point operations using an expanded-range floating-point format in processor devices is disclosed herein. In some aspects, a processor device comprises a floating-point unit (FPU) that is configured to perform a floating-point operation using one or more floating-point values interpreted by the FPU according to an expanded-range floating-point format that comprises no sign bit, a plurality of exponent bits including an additional exponent bit relative to a corresponding standard floating-point format, and a plurality of significand bits. The expanded-range floating-point format comprises a same number of bits as the corresponding standard floating-point format. A count of the plurality of significand bits of the expanded-range floating-point format is identical to a count of a plurality of significand bits of the corresponding standard floating-point format, and the significand bits of the expanded-range floating-point format are interpreted in a manner identical to the significand bits of the corresponding standard floating-point format.
Owner:QUALCOMM INC

Floating point unit exception handling method and device of industrial safety operating system and medium

The invention provides a floating point unit exception handling method and device of an industrial safety operating system and a medium. The method comprises the following steps: hooking a floating point exception handling function at an undefined instruction exception entry, and enabling the exception of a floating point unit to enter a unified handling flow; reading an abnormal comprehensive register of the processor to identify the floating point operation exception type, and reading a program counter register to obtain a trigger instruction address; performing standardized calculation on the operation under abnormal conditions according to a binary floating point arithmetic standard, generating result data corresponding to the operation, and forming abnormal mark information; writing the result data back to the destination register, and calculating a next instruction address according to the instruction set rule; and outputting the trigger instruction address, the exception type, the source operand and the result data as exception information through an interface from the kernel to a user side. According to the method, instant processing of the abnormal trigger point, continuous execution of the instruction, real-time output of the abnormal context and reduction of positioning and recovery time delay can be realized.
Owner:BEIJING HOLLYSYS TECHNOLOGY RESEARCH INSTITUTE CO LTD

System and Method for Determining Compression Performance of Codebooks Using Integer-Only Calculations for Resource-Constrained Environments

PendingUS20250284394A1Input/output to record carriersCode conversionMicrocontrollerFloating-point unit
A system and methods for determining compression performance of codebooks in resource-constrained computing environments using only integer operations. The system enables accurate prediction of compression performance without codebook generation on devices lacking floating-point units by utilizing bit manipulation techniques. Sourceblock occurrences are tracked using integer counters, with sum of squared probabilities calculated through integer multiplication. Logarithmic approximations are performed using most significant bit position analysis and bit-shifts replace division operations when sourceblock lengths are constrained to powers of 2. This approach requires minimal memory resources, just four integer registers beyond occurrence counters, making compression optimization viable on ultra-low-power 32-bit microcontrollers with less than 64 KB memory, such as those used in wearables and IoT devices, while maintaining accuracy within one bit of theoretical optimal performance.
Owner:ATOMBEAM TECH INC

Computational Optimization of Low-Precision Machine Learning Operations

ActiveCN112330523BResource allocationDigital data processing detailsGraphicsGeneral purpose graphical processing unit
The title of the present invention is "Computational Optimization for Low-Precision Machine Learning Operations". One embodiment provides a general-purpose graphics processing unit including a dynamic precision floating-point unit, the dynamic precision floating-point unit including a control unit having precision tracking hardware logic to track the available number of precision bits of computational data related to a target precision, wherein the dynamic precision floating-point unit includes computational logic to output data at multiple precisions.
Owner:INTEL CORP

Fully configurable floating-point format

One embodiment provides a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets connected to the plurality of chiplet sockets. At least one of the plurality of chiplets comprises a graphics processing cluster including a plurality of processing resources. At least one processing resource of the plurality of processing resources includes a dynamic precision floating-point unit having floating-point circuitry configured to process an input in a configurable floating-point format with a variable number of exponent and mantissa bits.
Owner:INTEL CORP

Compute optimizations for low precision machine learning operations

One embodiment provides a general-purpose graphics processing unit comprising a dynamic precision floating-point unit including a control unit having precision tracking hardware logic to track an available number of bits of precision for computed data relative to a target precision, wherein the dynamic precision floating-point unit includes computational logic to output data at multiple precisions.
Owner:INTEL CORP

Register file for systolic array

A processing apparatus includes a general-purpose parallel processing engine including a set of multiple processing elements including a single precision floating-point unit, a double precision floating point unit, and an integer unit; a matrix accelerator including one or more systolic arrays; a first register file coupled with a first read control circuit, wherein the first read control circuit couples with the set of multiple processing elements and the matrix accelerator to arbitrate read requests to the first register file from the set of multiple processing elements and the matrix accelerator; and a second register file coupled with a second read control circuit, wherein the second read control circuit couples with the matrix accelerator to arbitrate read requests to the second register file from the matrix accelerator and limit access to the second register file by the set of multiple processing elements.
Owner:INTEL CORP

Computational optimization for low-precision machine learning operations

One embodiment provides a general purpose graphics processing unit including a dynamic precision floating point unit, the dynamic precision floating point unit including a control unit having precision tracking hardware logic to track the available number of precision bits of computation data relative to a target precision, wherein the dynamic precision floating point unit includes computation logic to output data at multiple precisions.
Owner:INTEL CORP

A low power multi-core shared floating point unit structure

This invention discloses a low-power multi-core shared floating-point unit structure, relating to the field of digital signal processing technology. It addresses the technical problem that existing floating-point units occupy a large silicon wafer area, leading to increased manufacturing costs and reduced yields after configuring one floating-point unit for each core. The low-power multi-core shared floating-point unit structure of this invention includes a processor core, shared floating-point units, and a multi-level logarithmic interconnect network. The processor core and the shared floating-point units communicate via an auxiliary processing unit interface designed for tightly coupled accelerators. The processor core executes application streams and dispatches floating-point calculation instructions to the shared floating-point unit cluster. The shared floating-point units perform floating-point operations as a shared computing resource time-division multiplexed by multiple processor cores. The multi-level logarithmic interconnect network dynamically and scalably connects all processor cores and shared floating-point units to form a communication path.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Computational optimization of low precision machine learning operations

The title of the invention is computational optimization of low precision machine learning operations. One embodiment provides a general purpose graphics processing unit including a dynamic precision floating point unit including a control unit having precision tracking hardware logic to track an available number of precision bits of computational data related to target precision, wherein the dynamic precision floating point unit includes computational logic to output data at a plurality of precisions.
Owner:INTEL CORP

Method for performing encoding format conversion with integer arithmetic operations and system therefor

A method for performing encoding format conversion with integer arithmetic operations and a system performing the method are provided. The system includes a target device that adopts a power-management bus (PMBus), and operates an embedded system that is only capable of integer arithmetic operation without any floating-point unit. In the method, a floating-point arithmetic encoding format value is inputted. After a logarithm value of the floating-point value is obtained by extracting an exponent value from a binary representation of the floating-point value, a new exponent value can be obtained through a control flow. A new mantissa value is then calculated according to the floating-point value, the exponent value and the new exponent value. A new value that is a linear encoding format value converted from the floating-point arithmetic encoding format value is generated by combining the new exponent value and the new mantissa value.
Owner:ADATA TECHNOLOGY CO LTD

System and Method for Integer-Only Implementation of Real-Time Codebook Performance Tracking for Resource-Constrained Environments

A system and methods for implementing real-time tracking of codebook compression performance using only integer operations in resource-constrained computing environments. The system transforms floating-point calculations into equivalent operations using only integer arithmetic, bit shifts, and bit manipulations, enabling deployment on ultra-low-power microcontrollers lacking floating-point units. By constraining calculations to additions, subtractions, multiplications, and bit shifts, the system maintains accurate performance tracking while dramatically reducing computational requirements and power consumption. The method normalizes parameters across different sourceblock lengths, implements logarithmic approximations using most significant bit (MSB) detection, and replaces divisions with bit shifts where possible. This approach makes sophisticated compression performance tracking viable on microcontroller processors and similar resource-constrained platforms, extending advanced data compression capabilities to billions of edge devices where energy efficiency is paramount.
Owner:ATOMBEAM TECH INC

System and Method for Generating Full Binary Tree Codebooks Using Only Integer Operations for Resource-Constrained Environments

PendingUS20250291487A1Input/output to record carriersCode conversionMicrocontrollerAlgorithm transformation
A system and methods for generating full binary tree codebooks for encoding data within one bit of the optimal expected word length using only integer operations. The system transforms the modified Shannon-Fano algorithm into an implementation requiring only additions, subtractions, multiplications, and bit shifts, eliminating floating-point operations entirely. By tracking occurrence counts directly, replacing logarithmic calculations with most significant bit position detection, and using bit shifts instead of division, the method enables codebook generation on ultra-low-power microcontrollers lacking floating-point units. The approach requires only four integer registers beyond the occurrence counters, maintains a full binary tree structure ensuring every bit pattern of a given length is a valid codeword, and achieves compression performance within one bit of optimal. This implementation extends advanced compression capabilities to billions of resource-constrained devices including IoT sensors, wearables, and embedded systems where power consumption and computational capacity are severely limited.
Owner:ATOMBEAM TECH INC

Performing floating-point operations using an expanded-range floating-point format in processor devices

Performing floating-point operations using an expanded-range floating-point format in processor devices is disclosed herein. In some aspects, a processor device comprises a floating-point unit (FPU) that is configured to perform a floating-point operation using one or more floating-point values interpreted by the FPU according to an expanded-range floating-point format that comprises no sign bit, a plurality of exponent bits including an additional exponent bit relative to a corresponding standard floating-point format, and a plurality of significand bits. The expanded-range floating-point format comprises a same number of bits as the corresponding standard floating-point format. A count of the plurality of significand bits of the expanded-range floating-point format is identical to a count of a plurality of significand bits of the corresponding standard floating-point format, and the significand bits of the expanded-range floating-point format are interpreted in a manner identical to the significand bits of the corresponding standard floating-point format.
Owner:QUALCOMM INC

System and Method for Integer-Only Hybrid Codebook Performance Estimation for Resource-Constrained Environments

A system and methods for determining compression performance of hybrid codebook systems using only integer operations for ultra-low-power devices. The system transforms the calculation of combined performance metrics for primary / secondary codebook pairs into equivalent operations using only additions, subtractions, multiplications, and bit shifts. By normalizing occurrence counters, transforming logarithmic calculations into MSB-based bit manipulations, and replacing divisions with bit shifts, the system enables accurate performance estimation without floating-point operations. This approach makes hybrid codebook optimization viable on microcontrollers without floating-point units, such as ARM Cortex-M series processors. The method maintains performance accuracy while dramatically reducing computational complexity and power consumption, enabling advanced compression techniques on resource-constrained IoT devices, wearables, and embedded systems where both compression efficiency and energy conservation are critical requirements.
Owner:ATOMBEAM TECH INC

Fully configurable floating-point format

One embodiment provides a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets coupled with the plurality of chiplet sockets. At least one of the plurality of chiplets comprising a graphics processing cluster including a plurality of processing resources. At least one processing resource of the plurality of processing resources includes a dynamic precision floating-point unit having floating-point circuitry configured to process input in a configurable floating-point format having a variable number of exponent and mantissa bits.
Owner:INTEL CORP

Fully configurable floating point format

The invention relates to a fully configurable floating point format. One embodiment provides a graphics processor comprising: a base die comprising a plurality of core grain sockets; and a plurality of core particles coupled with the plurality of core particle sockets. At least one of the plurality of cores includes a graphics processing cluster including a plurality of processing resources. At least one of the plurality of processing resources includes a dynamic precision floating point unit having a floating point circuit configured to process an input in a configurable floating point format having a variable number of exponent bits and mantissa bits.
Owner:INTEL CORP

STORING FLOATING-POINT VALUES ACCORDING TO AN EXTENDED QFLOAT FLOATING-POINT (xqFP) FORMAT IN PROCESSOR DEVICES

PCT designated stage expiredWO2025155499A1Digital data processing detailsSign bitSoftware engineering
Storing floating-point values according to an extended QFloat floating-point (xqFP) format in processor devices is disclosed herein. In some aspects, a processor device comprises a register file comprising a plurality of registers, and comprises a floating-point unit (FPU) circuit that is configured to store a first floating-point value in a register of the plurality of registers. The first floating-point value is formatted according to the xqFP format that comprises an exponent field and a significand field. The significand field is formatted as a signed one's complement value, and comprises a sign bit, an explicit most-significant-bit (MSB), a fractional field, and a deferred increment bit that represents a value of one-half (½) unit of least precision (ULP).
Owner:QUALCOMM INC

Image processing method, readable medium and electronic device

The application relates to the technical field of artificial intelligence, and discloses an image processing method, a readable medium and an electronic device. The method comprises the following steps: acquiring a to-be-recognized image feature map; acquiring floating-point region position information of a region of interest on a to-be-recognized image, and quantizing the floating-point unit position information into fixed-point region position information; based on the fixed-point region position information of the region of interest, floating-point unit position information of a region of interest unit included in the region of interest is obtained; the floating-point unit position information is quantized to obtain fixed-point unit position information; based on the fixed-point unit position information, the feature values of the corresponding region of interest unit are acquired; based on the feature values of the region of interest unit in the region of interest, a region feature map of the region of interest is obtained, and a recognition result of the to-be-recognized image is obtained. By quantizing the position information of the region of interest in the to-be-recognized image and the position information of each unit of the region of interest, the running speed of a neural network model can be improved.
Owner:ARM TECH CHINA CO LTD

Method, medium, and apparatus for computational optimization of low-precision machine learning operations

ActiveCN116414455BGeneral purposeGraphics
The present application is entitled "Computational Optimization of Low Precision Machine Learning Operations". One embodiment provides a general purpose graphics processing unit including a dynamic precision floating point unit including a control unit having precision tracking hardware logic to track an available number of precision bits of computational data related to a target precision, wherein the dynamic precision floating point unit includes computational logic to output data at a plurality of precisions.
Owner:INTEL CORP

Differential pipeline delays in a coprocessor

A coprocessor such as a floating-point unit includes a pipeline that is partitioned into a first portion and a second portion. A controller is configured to provide control signals to the first portion and the second portion of the pipeline. A first physical distance traversed by control signals propagating from the controller to the first portion of the pipeline is shorter than a second physical distance traversed by control signals propagating from the controller to the second portion of the pipeline. A scheduler is configured to cause a physical register file to provide a first subset of bits of an instruction to the first portion at a first time. The physical register file provides a second subset of the bits of the instruction to the second portion at a second time subsequent to the first time.
Owner:ADVANCED MICRO DEVICES INC

IoT (Internet of Things) sensing layer lightweight encryption method and system based on PUF (Physical Unclonable Function) and chaotic key generation

The invention discloses an IoT (Internet of Things) sensing layer lightweight encryption method and system based on PUF (Physical Unclonable Function) and chaotic key generation, and relates to the technical field of Internet of Things security. When equipment is powered on, an SRAM (Static Random Access Memory) PUF is utilized to generate a unique and unclonable hardware random source, a stable chaotic seed sequence is recovered in combination with BCH (Broadcast Channel) fuzzy extraction, Helper Data which does not contain sensitive information is subjected to on-chain index management, and the Helper Data is subjected to on-chain index management; realizing cross-device consistency of key parameters and device identity binding; in addition, a dynamic chaos parameter updating and exception handling mechanism is constructed, when parameter failure or side channel attack risk is detected, PUF depth resampling is triggered, and safety parameters are synchronously recovered through on-chain increments. A floating point unit is not needed in the whole process, and the high-safety and low-power-consumption real-time encryption requirements of resource-limited IoT sensing layer equipment such as medical treatment, physiological monitoring and supply chains can be met.
Owner:NANJING UNIVERSTIY SUZHOU HIGH TECH INST

Lightweight real-time target tracking system and method applied to embedded microcontroller

The invention discloses a lightweight target tracking method and system for a resource-constrained embedded platform, and mainly solves the problem that a complex target tracking algorithm is difficult to give consideration to real-time performance and robustness on a low-cost microcontroller. The core of the invention lies in providing an improved tracking strategy fused with inertial navigation information. The method specifically comprises the steps of introducing IMU inertial navigation data to construct a feed-forward motion compensation mechanism aiming at the problem of view field offset caused by attitude change of an observation platform in an air-to-air tracking scene, calculating pixel displacement caused by attitude change of a carrier platform in real time, and removing the pixel displacement in a Kalman filtering prediction stage to realize decoupling of target motion and platform motion. According to the method, deep appearance feature extraction is abandoned, and a self-adaptive Kalman filtering algorithm is provided. According to the algorithm, measurement noise covariance is dynamically adjusted in real time by using detection confidence output by a front end, so that track oscillation caused by low-quality observation is inhibited; and meanwhile, a process noise covariance is adaptively increased by using normalized information square (NIS) statistical characteristics, and tracking lag under fast maneuvering (jitter) is eliminated. A static memory pool technology is adopted on a microcontroller to manage a track life cycle, and a hardware floating point unit (FPU) and a CMSIS-DSP math library are utilized to accelerate matrix operation. According to the invention, low-delay and high-precision target tracking is realized on the STM32 microcontroller, and the method is widely applied to an embedded microcontroller platform.
Owner:NANJING UNIV OF SCI & TECH

Bitwidth reconfiguration of register files using shadow latch configuration

A processor comprising a front end having an instruction set, the front end operating at a first bit width; and a floating point unit coupled to receive the instruction set in the processor operating at the first bit width. The floating point unit operates at a second bit width, and based on a bit width evaluation of the instruction set provided to the floating point unit, the floating point unit employs a shadow latch configured floating point unit register file to perform a bit width reconfiguration. The shadow latch configured floating point register file includes a plurality of regular latches and a plurality of shadow latches for storing data to be read from or written to the shadow latches. The bit width reconfiguration enables the floating point unit operating at the second bit width to operate on the instruction set received at the first bit width.
Owner:ADVANCED MICRO DEVICES INC