1ULP Floating-Point Sum of Squares Hardware for Graphics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current graphics processors require significant silicon area and power for performing floating-point n-input sum of squares operations, which is a common operation in geometry computation and machine learning tasks, and existing solutions like IEEE rounding are not fully necessary for all applications.

Innovation Solution

Implementing an efficient floating-point n-input sum of squares operation using 1 unit in the place (ULP) hardware, which is smaller, faster, and more power-efficient, as many graphics and machine learning algorithms do not require fully correct rounding for every operation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If IEEE rounding hardware is used for floating-point n-input sum of squares operations, then calculation precision is improved, but silicon area and power consumption increase

Engineering Contradiction:
Improvecalculation precisionVSAvoidsilicon area
Core Design Contradiction:
Measurement precisionVSArea of stationary object

Solution Approach 1:

The patent changes the precision parameter from full IEEE rounding to 1ulp (unit in the last place) precision. This parameter change allows the hardware to achieve acceptable precision for graphics and machine learning applications while significantly reducing the silicon area required for the sum of squares calculation unit.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If IEEE rounding hardware is used for floating-point n-input sum of squares operations, then calculation precision is improved, but power consumption increases

Engineering Contradiction:
Improvecalculation precisionVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

The patent changes the precision parameter from full IEEE rounding to 1ulp precision, which reduces the computational complexity and power consumption of the hardware while maintaining sufficient accuracy for the intended applications in graphics processing and machine learning.

Inventive Principle:
Principle #35Parameter changes

3Area of stationary object

If 1ulp hardware is used for floating-point n-input sum of squares operations, then silicon area and power consumption are reduced, but calculation precision decreases

Engineering Contradiction:
Improvesilicon areaVSAvoidcalculation precision
Core Design Contradiction:
Area of stationary objectVSMeasurement precision

Solution Approach 1:

The patent applies partial action by implementing only 1ulp precision instead of full IEEE rounding. This partial implementation is sufficient for the required applications in graphics and machine learning, achieving the desired balance between precision and hardware efficiency without the excessive precision that full IEEE rounding would provide.

Inventive Principle:
Principle #16Partial or excessive action

4Use of energy by stationary object

If 1ulp hardware is used for floating-point n-input sum of squares operations, then power consumption is reduced, but calculation precision decreases

Engineering Contradiction:
Improvepower consumptionVSAvoidcalculation precision
Core Design Contradiction:
Use of energy by stationary objectVSMeasurement precision

Solution Approach 1:

The patent changes the precision parameter from full IEEE rounding to 1ulp precision, which reduces the computational complexity and power consumption of the hardware while maintaining sufficient accuracy for the intended applications in graphics processing and machine learning.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240143279A1Floating-point n-input sum of squares 1ULP hardware
Publication Date: 2024.05.02 INTEL CORP
  • US20240143279A1 patent drawing
  • US20240143279A1 patent drawing
  • US20240143279A1 patent drawing

AI summary

Described herein is a technique to implement an efficient floating-point n-input sum of squares operation using faithful rounding to 1 unit in the place (ULP) instead of IEEE rounding. The resulting circuitry is useful to accelerate graphics algorithms that don't require fully IEEE compliant hardware. Multipliers that are 1ulp can be significantly smaller, faster and more power efficient than IEEE rounded multipliers.