Multiplier-Accumulator Utilization for Mixed-Precision Floating-Point Operations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing design of multiplier-accumulators in logical operation units leads to low utilization rates and increased hardware overhead due to idle half-precision and single-precision multiplier-accumulators during floating-point number multiplication-accumulation operations, as they are not optimally utilized when performing operations of different precision levels.

Innovation Solution

A high-performance multiplier-accumulator is designed with N single-precision multiplication-accumulation units, each comprising two half-precision multiplier-accumulators, and an asymmetric multiplier-accumulator with N units combining single-precision and half-precision multiplier-accumulators, allowing for efficient operation across single-precision and half-precision floating-point number multiplication-accumulation tasks, thereby improving utilization and reducing hardware overhead.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If n single-precision multiplier-accumulators and 2n half-precision multiplier-accumulators are arranged in the logical operation unit, then the multiplication-accumulation operations for both single-precision and half-precision floating-point numbers can be performed simultaneously, but the utilization rate of the multiplier-accumulators decreases and hardware overhead increases

Engineering Contradiction:
Improvecapability to perform both single-precision and half-precision multiplication-accumulation operationsVSAvoidhardware overhead and utilization rate
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent makes half-precision multiplier-accumulators capable of performing both half-precision and single-precision multiplication-accumulation operations by introducing a conversion unit that converts single-precision floating-point numbers to half-precision format. This allows the same hardware resource (half-precision multiplier-accumulators) to serve multiple precision requirements, eliminating the need for separate single-precision units and improving overall utilization rate.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges the functionality of separate single-precision and half-precision multiplier-accumulators into a unified structure where half-precision units can handle both precision types. By combining the conversion unit with the half-precision multiplier-accumulators, the system reduces the total number of required units from 3n (n single-precision + 2n half-precision) to just 2n half-precision units, thereby reducing hardware overhead while maintaining full operational capability.

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If 2n half-precision multiplier-accumulators are used for single-precision operations, then single-precision multiplication-accumulation can be performed, but the half-precision multiplier-accumulators operate inefficiently with increased hardware overhead

Engineering Contradiction:
Improvesingle-precision multiplication-accumulation capabilityVSAvoidhardware overhead
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent changes the precision parameter of the input data by converting single-precision floating-point numbers to half-precision format before processing. This parameter transformation allows the half-precision multiplier-accumulators to operate in their native precision mode while still handling single-precision computational tasks, thereby improving efficiency and reducing the need for additional single-precision hardware units.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If n single-precision multiplier-accumulators are used for half-precision operations, then half-precision multiplication-accumulation can be performed, but the single-precision multiplier-accumulators are underutilized and hardware overhead increases

Engineering Contradiction:
Improvehalf-precision multiplication-accumulation capabilityVSAvoidhardware overhead and utilization rate
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent enables half-precision multiplier-accumulators to universally handle both half-precision and single-precision operations through data conversion. This multi-functionality eliminates the need for dedicated single-precision units, as the same half-precision units can be dynamically configured to handle different precision requirements, thereby improving utilization rates and reducing overall hardware overhead.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240020094A1Multiplication-accumulation system, multiplication-accumulation method, and electronic device
Publication Date: 2024.01.18 GLENFLY TECH CO LTD
  • US20240020094A1 patent drawing
  • US20240020094A1 patent drawing
  • US20240020094A1 patent drawing

AI summary

Multiplication-accumulation method and apparatus, a processor, and a computer program product are provided. The method includes: when a logical operation unit performs single-precision floating-point number multiplication-accumulation operation, combining two half-precision multiplier-accumulators in each single-precision multiplication-accumulation unit to perform the multiplication-accumulation operation on to-be-processed single-precision floating-point numbers to obtain corresponding single-precision multiplication-accumulation results, a total of N multiplication-accumulation results being obtained; and when the logical operation unit performs half-precision floating-point number multiplication-accumulation operation, performing, by each half-precision multiplier-accumulator, the multiplication-accumulation operation on to-be-processed half-precision floating-point numbers to obtain corresponding half-precision multiplication-accumulation results, a total of 2N multiplication-accumulation results being obtained. Utilization of the multiplier-accumulators is improved.