Fuzzy-Jbit Floating-Point Multiply-Accumulate for Single-Cycle Throughput

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer processor architectures face challenges in efficiently handling large matrices, particularly in deep learning applications, where matrix operations are instruction intensive and require significant resources, leading to performance bottlenecks in operations like matrix multiplication and floating-point arithmetic.

Innovation Solution

The introduction of a matrix operations accelerator that utilizes 2-dimensional data structures called 'tiles' for efficient matrix processing, along with the implementation of a Fuzzy-Jbit format for floating-point multiply-accumulate instructions, allowing for improved performance and energy efficiency in matrix operations by optimizing memory access and arithmetic operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional floating-point FMA instructions are used with strict format requirements, then precision is maintained, but hardware complexity and cycle time increase

Engineering Contradiction:
Improvefloating-point precisionVSAvoidhardware complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes the parameter of Jbit position from a fixed location (bit 23) to a variable location that can be at bit 23 or bit 24. This parameter change allows the hardware to accommodate different alignment scenarios without requiring complex end-of-cycle adjustment logic, thereby reducing hardware complexity while maintaining precision through controlled rounding modes.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces dynamic behavior by allowing the Jbit position to vary between cycles based on the alignment of mantissas. The system dynamically adjusts the Jbit position and uses different rounding modes (round-to-nearest-even or round-toward-zero) depending on the specific operation, which simplifies the hardware compared to enforcing a strict fixed Jbit position.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If strict floating-point format is enforced with fixed Jbit position, then precision is maintained, but operation time increases due to end-of-cycle adjustments

Engineering Contradiction:
Improvefloating-point precisionVSAvoidoperation cycle time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary alignment of mantissas before the FMA operation by adjusting the Jbit position in advance. This preliminary action ensures that the addition can proceed without requiring complex end-of-cycle adjustments, thereby reducing the operation cycle time while maintaining precision through proper rounding.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the Jbit position parameter dynamically based on the alignment needs of the operation. By allowing Jbit to be at different positions (23 or 24) and using appropriate rounding modes, the system eliminates the need for time-consuming end-of-cycle adjustments while preserving precision.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If conventional matrix multiplication methods are used, then accuracy is maintained, but throughput decreases due to instruction intensity

Engineering Contradiction:
Improvecalculation accuracyVSAvoidthroughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the FMA operation into distinct phases with variable Jbit positioning, allowing different parts of the computation to use optimized paths. This segmentation enables higher throughput by avoiding uniform complex processing for all operations, while maintaining accuracy through proper rounding in each segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes operational parameters (Jbit position and rounding mode) based on the specific computation needs, allowing high-throughput paths for operations that don't require strict adherence to traditional formatting, while maintaining accuracy for operations that do.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If traditional FMA hardware is implemented, then precision is maintained, but energy consumption increases

Engineering Contradiction:
Improvearithmetic precisionVSAvoidenergy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent changes the Jbit position parameter and rounding mode based on operational needs, allowing the hardware to use simpler, lower-energy paths for operations that don't require full precision adjustment, while maintaining precision when needed. This reduces overall energy consumption compared to always using the full precision adjustment path.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3716050B1Using fuzzy-jbit location of floating-point multiply-accumulate results
Publication Date: 2023.06.07 INTEL CORP
  • EP3716050B1 patent drawingFigure 1A~1B
  • EP3716050B1 patent drawingFigure 2(A)~2(C)
  • EP3716050B1 patent drawingFigure 3

AI summary

Disclosed embodiments relate to performing floating-point (FP) arithmetic. In one example, a processor is to decode an instruction specifying locations of first, second, and third floating-point (FP) operands and an opcode calling for accumulating a FP product of the first and second FP operands with the third FP operand, and execution circuitry to, in a first cycle, generate the FP product having a Fuzzy-Jbit format comprising a sign bit, a 9-bit exponent, and a 25-bit mantissa having two possible positions for a JBit and, in a second cycle, to accumulate the FP product with the third FP operand, while concurrently, based on Jbit positions of the FP product and the third FP operand, determining an exponent adjustment and a mantissa shift control of a result of the accumulation, wherein performing the exponent adjustment concurrently enhances an ability to perform the accumulation in one cycle.