Matrix Multiplication Hardware Selecting Adder Tree Intermediate Results

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current hardware solutions for matrix multiplication, particularly in artificial intelligence computations like convolution, face inefficiencies with small input and output channel sizes, leading to low utilization and output fragmentation, which increases hardware resource requirements.

Innovation Solution

A matrix multiplication hardware system with a hierarchical tree of adders and a control unit that allows for parallel performance of different matrix multiplications using a combined matrix, enabling the selection of intermediate results to optimize output and reduce data fragmentation, specifically designed to handle various size inputs and improve efficiency in groupwise convolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If hardware solutions are used to perform matrix multiplication, then computational performance is improved, but hardware resource requirements increase when dealing with small input and output channel sizes

Engineering Contradiction:
Improvecomputational performanceVSAvoidhardware resource requirements
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the matrix multiplication operation into multiple independent multiplication operations that can be performed in parallel. By dividing the computation into separate operations, the hardware can process multiple tasks simultaneously, improving productivity while managing hardware resource usage through efficient parallel execution rather than requiring a single complex hardware solution for all cases.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The hardware solution is designed to be universal and can perform different matrix multiplication operations with varying input and output channel sizes. The system can handle both small and large channel sizes using the same hardware architecture, eliminating the need for separate specialized hardware for different scenarios and reducing overall hardware resource requirements while maintaining high computational performance.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If conventional matrix multiplication hardware is used, then large channel sizes are handled efficiently, but small channel sizes result in low utilization and output fragmentation

Engineering Contradiction:
Improveefficiency for large channelsVSAvoidperformance consistency across different channel sizes
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements dynamic adaptation where the hardware can adjust its operation mode based on the input channel size. For small channel sizes, the system can dynamically reconfigure to process operations more efficiently, avoiding fixed inefficient patterns. This dynamic behavior allows the same hardware to maintain high efficiency across both small and large channel sizes, ensuring performance consistency without sacrificing reliability for any specific channel dimension.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes operational parameters such as the number and configuration of parallel multiplication operations based on the input channel size. By adjusting these parameters dynamically, the hardware can optimize its performance for both small and large channels, preventing the low utilization and output fragmentation that occur with fixed hardware configurations designed only for large channel sizes.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If hardware solutions are designed for specific operations, then performance is optimized for that operation, but flexibility to handle different matrix operations is reduced

Engineering Contradiction:
Improveoperation-specific performanceVSAvoidsupport for different matrix operations
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The hardware solution is designed with multi-functionality to support various matrix operations including convolution, groupwise convolution, and other linear algebra operations. By incorporating universal features that can handle different operation types, the system maintains high performance for each specific operation while also providing flexibility and adaptability across different computational tasks, eliminating the trade-off between specialization and versatility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements partial parallelization where not all operations need to be performed in parallel, but the hardware is designed to support selective parallel execution. This allows the system to optimize performance for specific operations when needed while maintaining the ability to handle different matrix operations with varying degrees of parallelism, thus achieving both high operation-specific performance and broad adaptability without requiring full parallelization for all cases.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11520854B2Support for different matrix multiplications by selecting adder tree intermediate results
Publication Date: 2022.12.06 META PLATFORMS INC
  • US11520854B2 patent drawing
  • US11520854B2 patent drawing
  • US11520854B2 patent drawing

AI summary

A first group of elements is element-wise multiplied with a second group of elements using a plurality of multipliers belonging to a matrix multiplication hardware unit. Results of the plurality of multipliers are added together using a hierarchical tree of adders belonging to the matrix multiplication hardware unit and a final result of the hierarchical tree of adders or any of a plurality of intermediate results of the hierarchical tree of adders is selectively provided for use in determining an output result matrix. A control unit is used to instruct the matrix multiplication hardware unit to perform a plurality of different matrix multiplications in parallel by using a combined matrix that includes elements of a plurality of different operand matrices and utilize one or more selected ones of the intermediate results of the hierarchical tree of adders for use in determining the output result matrix that includes different groups of elements representing different multiplication results corresponding to different ones of the different operand matrices.