Fuzzy-Jbit Floating-Point Multiply-Accumulate for Single-Cycle Throughput
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer processor architectures face challenges in efficiently handling large matrices, particularly in deep learning applications, where matrix operations are instruction intensive and require significant resources, leading to performance bottlenecks in operations like matrix multiplication and floating-point arithmetic.
Innovation Solution
The introduction of a matrix operations accelerator that utilizes 2-dimensional data structures called 'tiles' for efficient matrix processing, along with the implementation of a Fuzzy-Jbit format for floating-point multiply-accumulate instructions, allowing for improved performance and energy efficiency in matrix operations by optimizing memory access and arithmetic operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional floating-point FMA instructions are used with strict format requirements, then precision is maintained, but hardware complexity and cycle time increase
Solution Approach 1:
The patent changes the parameter of Jbit position from a fixed location (bit 23) to a variable location that can be at bit 23 or bit 24. This parameter change allows the hardware to accommodate different alignment scenarios without requiring complex end-of-cycle adjustment logic, thereby reducing hardware complexity while maintaining precision through controlled rounding modes.
Solution Approach 2:
The patent introduces dynamic behavior by allowing the Jbit position to vary between cycles based on the alignment of mantissas. The system dynamically adjusts the Jbit position and uses different rounding modes (round-to-nearest-even or round-toward-zero) depending on the specific operation, which simplifies the hardware compared to enforcing a strict fixed Jbit position.
2Measurement precision
If strict floating-point format is enforced with fixed Jbit position, then precision is maintained, but operation time increases due to end-of-cycle adjustments
Solution Approach 1:
The patent performs preliminary alignment of mantissas before the FMA operation by adjusting the Jbit position in advance. This preliminary action ensures that the addition can proceed without requiring complex end-of-cycle adjustments, thereby reducing the operation cycle time while maintaining precision through proper rounding.
Solution Approach 2:
The patent changes the Jbit position parameter dynamically based on the alignment needs of the operation. By allowing Jbit to be at different positions (23 or 24) and using appropriate rounding modes, the system eliminates the need for time-consuming end-of-cycle adjustments while preserving precision.
3Measurement precision
If conventional matrix multiplication methods are used, then accuracy is maintained, but throughput decreases due to instruction intensity
Solution Approach 1:
The patent segments the FMA operation into distinct phases with variable Jbit positioning, allowing different parts of the computation to use optimized paths. This segmentation enables higher throughput by avoiding uniform complex processing for all operations, while maintaining accuracy through proper rounding in each segment.
Solution Approach 2:
The patent changes operational parameters (Jbit position and rounding mode) based on the specific computation needs, allowing high-throughput paths for operations that don't require strict adherence to traditional formatting, while maintaining accuracy for operations that do.
4Measurement precision
If traditional FMA hardware is implemented, then precision is maintained, but energy consumption increases
Solution Approach 1:
The patent changes the Jbit position parameter and rounding mode based on operational needs, allowing the hardware to use simpler, lower-energy paths for operations that don't require full precision adjustment, while maintaining precision when needed. This reduces overall energy consumption compared to always using the full precision adjustment path.
Data Source
Figure 1A~1B
Figure 2(A)~2(C)
Figure 3
AI summary
Disclosed embodiments relate to performing floating-point (FP) arithmetic. In one example, a processor is to decode an instruction specifying locations of first, second, and third floating-point (FP) operands and an opcode calling for accumulating a FP product of the first and second FP operands with the third FP operand, and execution circuitry to, in a first cycle, generate the FP product having a Fuzzy-Jbit format comprising a sign bit, a 9-bit exponent, and a 25-bit mantissa having two possible positions for a JBit and, in a second cycle, to accumulate the FP product with the third FP operand, while concurrently, based on Jbit positions of the FP product and the third FP operand, determining an exponent adjustment and a mantissa shift control of a result of the accumulation, wherein performing the exponent adjustment concurrently enhances an ability to perform the accumulation in one cycle.