Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

15 results about "Floating-point unit" patented technology

A floating-point unit (FPU, colloquially a math coprocessor) is a part of a computer system specially designed to carry out operations on floating point numbers. Typical operations are addition, subtraction, multiplication, division, square root, and bitshifting. Some systems (particularly older, microcode-based architectures) can also perform various transcendental functions such as exponential or trigonometric calculations, though in most modern processors these are done with software library routines.

Floating point unit exception handling method and device of industrial safety operating system and medium

The invention provides a floating point unit exception handling method and device of an industrial safety operating system and a medium. The method comprises the following steps: hooking a floating point exception handling function at an undefined instruction exception entry, and enabling the exception of a floating point unit to enter a unified handling flow; reading an abnormal comprehensive register of the processor to identify the floating point operation exception type, and reading a program counter register to obtain a trigger instruction address; performing standardized calculation on the operation under abnormal conditions according to a binary floating point arithmetic standard, generating result data corresponding to the operation, and forming abnormal mark information; writing the result data back to the destination register, and calculating a next instruction address according to the instruction set rule; and outputting the trigger instruction address, the exception type, the source operand and the result data as exception information through an interface from the kernel to a user side. According to the method, instant processing of the abnormal trigger point, continuous execution of the instruction, real-time output of the abnormal context and reduction of positioning and recovery time delay can be realized.
Owner:BEIJING HOLLYSYS TECHNOLOGY RESEARCH INSTITUTE CO LTD

Fully configurable floating-point format

One embodiment provides a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets connected to the plurality of chiplet sockets. At least one of the plurality of chiplets comprises a graphics processing cluster including a plurality of processing resources. At least one processing resource of the plurality of processing resources includes a dynamic precision floating-point unit having floating-point circuitry configured to process an input in a configurable floating-point format with a variable number of exponent and mantissa bits.
Owner:INTEL CORP

Register file for systolic array

A processing apparatus includes a general-purpose parallel processing engine including a set of multiple processing elements including a single precision floating-point unit, a double precision floating point unit, and an integer unit; a matrix accelerator including one or more systolic arrays; a first register file coupled with a first read control circuit, wherein the first read control circuit couples with the set of multiple processing elements and the matrix accelerator to arbitrate read requests to the first register file from the set of multiple processing elements and the matrix accelerator; and a second register file coupled with a second read control circuit, wherein the second read control circuit couples with the matrix accelerator to arbitrate read requests to the second register file from the matrix accelerator and limit access to the second register file by the set of multiple processing elements.
Owner:INTEL CORP

Computational optimization for low-precision machine learning operations

One embodiment provides a general purpose graphics processing unit including a dynamic precision floating point unit, the dynamic precision floating point unit including a control unit having precision tracking hardware logic to track the available number of precision bits of computation data relative to a target precision, wherein the dynamic precision floating point unit includes computation logic to output data at multiple precisions.
Owner:INTEL CORP

A low power multi-core shared floating point unit structure

This invention discloses a low-power multi-core shared floating-point unit structure, relating to the field of digital signal processing technology. It addresses the technical problem that existing floating-point units occupy a large silicon wafer area, leading to increased manufacturing costs and reduced yields after configuring one floating-point unit for each core. The low-power multi-core shared floating-point unit structure of this invention includes a processor core, shared floating-point units, and a multi-level logarithmic interconnect network. The processor core and the shared floating-point units communicate via an auxiliary processing unit interface designed for tightly coupled accelerators. The processor core executes application streams and dispatches floating-point calculation instructions to the shared floating-point unit cluster. The shared floating-point units perform floating-point operations as a shared computing resource time-division multiplexed by multiple processor cores. The multi-level logarithmic interconnect network dynamically and scalably connects all processor cores and shared floating-point units to form a communication path.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

System and Method for Generating Full Binary Tree Codebooks Using Only Integer Operations for Resource-Constrained Environments

PendingUS20250291487A1Input/output to record carriersCode conversionMicrocontrollerAlgorithm transformation
A system and methods for generating full binary tree codebooks for encoding data within one bit of the optimal expected word length using only integer operations. The system transforms the modified Shannon-Fano algorithm into an implementation requiring only additions, subtractions, multiplications, and bit shifts, eliminating floating-point operations entirely. By tracking occurrence counts directly, replacing logarithmic calculations with most significant bit position detection, and using bit shifts instead of division, the method enables codebook generation on ultra-low-power microcontrollers lacking floating-point units. The approach requires only four integer registers beyond the occurrence counters, maintains a full binary tree structure ensuring every bit pattern of a given length is a valid codeword, and achieves compression performance within one bit of optimal. This implementation extends advanced compression capabilities to billions of resource-constrained devices including IoT sensors, wearables, and embedded systems where power consumption and computational capacity are severely limited.
Owner:ATOMBEAM TECH INC

System and Method for Integer-Only Hybrid Codebook Performance Estimation for Resource-Constrained Environments

A system and methods for determining compression performance of hybrid codebook systems using only integer operations for ultra-low-power devices. The system transforms the calculation of combined performance metrics for primary / secondary codebook pairs into equivalent operations using only additions, subtractions, multiplications, and bit shifts. By normalizing occurrence counters, transforming logarithmic calculations into MSB-based bit manipulations, and replacing divisions with bit shifts, the system enables accurate performance estimation without floating-point operations. This approach makes hybrid codebook optimization viable on microcontrollers without floating-point units, such as ARM Cortex-M series processors. The method maintains performance accuracy while dramatically reducing computational complexity and power consumption, enabling advanced compression techniques on resource-constrained IoT devices, wearables, and embedded systems where both compression efficiency and energy conservation are critical requirements.
Owner:ATOMBEAM TECH INC

Fully configurable floating-point format

One embodiment provides a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets coupled with the plurality of chiplet sockets. At least one of the plurality of chiplets comprising a graphics processing cluster including a plurality of processing resources. At least one processing resource of the plurality of processing resources includes a dynamic precision floating-point unit having floating-point circuitry configured to process input in a configurable floating-point format having a variable number of exponent and mantissa bits.
Owner:INTEL CORP

Fully configurable floating point format

The invention relates to a fully configurable floating point format. One embodiment provides a graphics processor comprising: a base die comprising a plurality of core grain sockets; and a plurality of core particles coupled with the plurality of core particle sockets. At least one of the plurality of cores includes a graphics processing cluster including a plurality of processing resources. At least one of the plurality of processing resources includes a dynamic precision floating point unit having a floating point circuit configured to process an input in a configurable floating point format having a variable number of exponent bits and mantissa bits.
Owner:INTEL CORP

Image processing method, readable medium and electronic device

The application relates to the technical field of artificial intelligence, and discloses an image processing method, a readable medium and an electronic device. The method comprises the following steps: acquiring a to-be-recognized image feature map; acquiring floating-point region position information of a region of interest on a to-be-recognized image, and quantizing the floating-point unit position information into fixed-point region position information; based on the fixed-point region position information of the region of interest, floating-point unit position information of a region of interest unit included in the region of interest is obtained; the floating-point unit position information is quantized to obtain fixed-point unit position information; based on the fixed-point unit position information, the feature values of the corresponding region of interest unit are acquired; based on the feature values of the region of interest unit in the region of interest, a region feature map of the region of interest is obtained, and a recognition result of the to-be-recognized image is obtained. By quantizing the position information of the region of interest in the to-be-recognized image and the position information of each unit of the region of interest, the running speed of a neural network model can be improved.
Owner:ARM TECH CHINA CO LTD

Method, medium, and apparatus for computational optimization of low-precision machine learning operations

ActiveCN116414455BGeneral purposeGraphics
The present application is entitled "Computational Optimization of Low Precision Machine Learning Operations". One embodiment provides a general purpose graphics processing unit including a dynamic precision floating point unit including a control unit having precision tracking hardware logic to track an available number of precision bits of computational data related to a target precision, wherein the dynamic precision floating point unit includes computational logic to output data at a plurality of precisions.
Owner:INTEL CORP

IoT (Internet of Things) sensing layer lightweight encryption method and system based on PUF (Physical Unclonable Function) and chaotic key generation

The invention discloses an IoT (Internet of Things) sensing layer lightweight encryption method and system based on PUF (Physical Unclonable Function) and chaotic key generation, and relates to the technical field of Internet of Things security. When equipment is powered on, an SRAM (Static Random Access Memory) PUF is utilized to generate a unique and unclonable hardware random source, a stable chaotic seed sequence is recovered in combination with BCH (Broadcast Channel) fuzzy extraction, Helper Data which does not contain sensitive information is subjected to on-chain index management, and the Helper Data is subjected to on-chain index management; realizing cross-device consistency of key parameters and device identity binding; in addition, a dynamic chaos parameter updating and exception handling mechanism is constructed, when parameter failure or side channel attack risk is detected, PUF depth resampling is triggered, and safety parameters are synchronously recovered through on-chain increments. A floating point unit is not needed in the whole process, and the high-safety and low-power-consumption real-time encryption requirements of resource-limited IoT sensing layer equipment such as medical treatment, physiological monitoring and supply chains can be met.
Owner:NANJING UNIVERSTIY SUZHOU HIGH TECH INST

Lightweight real-time target tracking system and method applied to embedded microcontroller

The invention discloses a lightweight target tracking method and system for a resource-constrained embedded platform, and mainly solves the problem that a complex target tracking algorithm is difficult to give consideration to real-time performance and robustness on a low-cost microcontroller. The core of the invention lies in providing an improved tracking strategy fused with inertial navigation information. The method specifically comprises the steps of introducing IMU inertial navigation data to construct a feed-forward motion compensation mechanism aiming at the problem of view field offset caused by attitude change of an observation platform in an air-to-air tracking scene, calculating pixel displacement caused by attitude change of a carrier platform in real time, and removing the pixel displacement in a Kalman filtering prediction stage to realize decoupling of target motion and platform motion. According to the method, deep appearance feature extraction is abandoned, and a self-adaptive Kalman filtering algorithm is provided. According to the algorithm, measurement noise covariance is dynamically adjusted in real time by using detection confidence output by a front end, so that track oscillation caused by low-quality observation is inhibited; and meanwhile, a process noise covariance is adaptively increased by using normalized information square (NIS) statistical characteristics, and tracking lag under fast maneuvering (jitter) is eliminated. A static memory pool technology is adopted on a microcontroller to manage a track life cycle, and a hardware floating point unit (FPU) and a CMSIS-DSP math library are utilized to accelerate matrix operation. According to the invention, low-delay and high-precision target tracking is realized on the STM32 microcontroller, and the method is widely applied to an embedded microcontroller platform.
Owner:NANJING UNIV OF SCI & TECH

Bitwidth reconfiguration of register files using shadow latch configuration

A processor comprising a front end having an instruction set, the front end operating at a first bit width; and a floating point unit coupled to receive the instruction set in the processor operating at the first bit width. The floating point unit operates at a second bit width, and based on a bit width evaluation of the instruction set provided to the floating point unit, the floating point unit employs a shadow latch configured floating point unit register file to perform a bit width reconfiguration. The shadow latch configured floating point register file includes a plurality of regular latches and a plurality of shadow latches for storing data to be read from or written to the shadow latches. The bit width reconfiguration enables the floating point unit operating at the second bit width to operate on the instruction set received at the first bit width.
Owner:ADVANCED MICRO DEVICES INC