Multiplier-Accumulator Utilization for Mixed-Precision Floating-Point Operations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing design of multiplier-accumulators in logical operation units leads to low utilization rates and increased hardware overhead due to idle half-precision and single-precision multiplier-accumulators during floating-point number multiplication-accumulation operations, as they are not optimally utilized when performing operations of different precision levels.
Innovation Solution
A high-performance multiplier-accumulator is designed with N single-precision multiplication-accumulation units, each comprising two half-precision multiplier-accumulators, and an asymmetric multiplier-accumulator with N units combining single-precision and half-precision multiplier-accumulators, allowing for efficient operation across single-precision and half-precision floating-point number multiplication-accumulation tasks, thereby improving utilization and reducing hardware overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If n single-precision multiplier-accumulators and 2n half-precision multiplier-accumulators are arranged in the logical operation unit, then the multiplication-accumulation operations for both single-precision and half-precision floating-point numbers can be performed simultaneously, but the utilization rate of the multiplier-accumulators decreases and hardware overhead increases
Solution Approach 1:
The patent makes half-precision multiplier-accumulators capable of performing both half-precision and single-precision multiplication-accumulation operations by introducing a conversion unit that converts single-precision floating-point numbers to half-precision format. This allows the same hardware resource (half-precision multiplier-accumulators) to serve multiple precision requirements, eliminating the need for separate single-precision units and improving overall utilization rate.
Solution Approach 2:
The patent merges the functionality of separate single-precision and half-precision multiplier-accumulators into a unified structure where half-precision units can handle both precision types. By combining the conversion unit with the half-precision multiplier-accumulators, the system reduces the total number of required units from 3n (n single-precision + 2n half-precision) to just 2n half-precision units, thereby reducing hardware overhead while maintaining full operational capability.
2Productivity
If 2n half-precision multiplier-accumulators are used for single-precision operations, then single-precision multiplication-accumulation can be performed, but the half-precision multiplier-accumulators operate inefficiently with increased hardware overhead
Solution Approach 1:
The patent changes the precision parameter of the input data by converting single-precision floating-point numbers to half-precision format before processing. This parameter transformation allows the half-precision multiplier-accumulators to operate in their native precision mode while still handling single-precision computational tasks, thereby improving efficiency and reducing the need for additional single-precision hardware units.
3Productivity
If n single-precision multiplier-accumulators are used for half-precision operations, then half-precision multiplication-accumulation can be performed, but the single-precision multiplier-accumulators are underutilized and hardware overhead increases
Solution Approach 1:
The patent enables half-precision multiplier-accumulators to universally handle both half-precision and single-precision operations through data conversion. This multi-functionality eliminates the need for dedicated single-precision units, as the same half-precision units can be dynamically configured to handle different precision requirements, thereby improving utilization rates and reducing overall hardware overhead.
Data Source
AI summary
Multiplication-accumulation method and apparatus, a processor, and a computer program product are provided. The method includes: when a logical operation unit performs single-precision floating-point number multiplication-accumulation operation, combining two half-precision multiplier-accumulators in each single-precision multiplication-accumulation unit to perform the multiplication-accumulation operation on to-be-processed single-precision floating-point numbers to obtain corresponding single-precision multiplication-accumulation results, a total of N multiplication-accumulation results being obtained; and when the logical operation unit performs half-precision floating-point number multiplication-accumulation operation, performing, by each half-precision multiplier-accumulator, the multiplication-accumulation operation on to-be-processed half-precision floating-point numbers to obtain corresponding half-precision multiplication-accumulation results, a total of 2N multiplication-accumulation results being obtained. Utilization of the multiplier-accumulators is improved.


