Matrix Multiplication Hardware Selecting Adder Tree Intermediate Results
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current hardware solutions for matrix multiplication, particularly in artificial intelligence computations like convolution, face inefficiencies with small input and output channel sizes, leading to low utilization and output fragmentation, which increases hardware resource requirements.
Innovation Solution
A matrix multiplication hardware system with a hierarchical tree of adders and a control unit that allows for parallel performance of different matrix multiplications using a combined matrix, enabling the selection of intermediate results to optimize output and reduce data fragmentation, specifically designed to handle various size inputs and improve efficiency in groupwise convolution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If hardware solutions are used to perform matrix multiplication, then computational performance is improved, but hardware resource requirements increase when dealing with small input and output channel sizes
Solution Approach 1:
The patent segments the matrix multiplication operation into multiple independent multiplication operations that can be performed in parallel. By dividing the computation into separate operations, the hardware can process multiple tasks simultaneously, improving productivity while managing hardware resource usage through efficient parallel execution rather than requiring a single complex hardware solution for all cases.
Solution Approach 2:
The hardware solution is designed to be universal and can perform different matrix multiplication operations with varying input and output channel sizes. The system can handle both small and large channel sizes using the same hardware architecture, eliminating the need for separate specialized hardware for different scenarios and reducing overall hardware resource requirements while maintaining high computational performance.
2Productivity
If conventional matrix multiplication hardware is used, then large channel sizes are handled efficiently, but small channel sizes result in low utilization and output fragmentation
Solution Approach 1:
The patent implements dynamic adaptation where the hardware can adjust its operation mode based on the input channel size. For small channel sizes, the system can dynamically reconfigure to process operations more efficiently, avoiding fixed inefficient patterns. This dynamic behavior allows the same hardware to maintain high efficiency across both small and large channel sizes, ensuring performance consistency without sacrificing reliability for any specific channel dimension.
Solution Approach 2:
The system changes operational parameters such as the number and configuration of parallel multiplication operations based on the input channel size. By adjusting these parameters dynamically, the hardware can optimize its performance for both small and large channels, preventing the low utilization and output fragmentation that occur with fixed hardware configurations designed only for large channel sizes.
3Productivity
If hardware solutions are designed for specific operations, then performance is optimized for that operation, but flexibility to handle different matrix operations is reduced
Solution Approach 1:
The hardware solution is designed with multi-functionality to support various matrix operations including convolution, groupwise convolution, and other linear algebra operations. By incorporating universal features that can handle different operation types, the system maintains high performance for each specific operation while also providing flexibility and adaptability across different computational tasks, eliminating the trade-off between specialization and versatility.
Solution Approach 2:
The patent implements partial parallelization where not all operations need to be performed in parallel, but the hardware is designed to support selective parallel execution. This allows the system to optimize performance for specific operations when needed while maintaining the ability to handle different matrix operations with varying degrees of parallelism, thus achieving both high operation-specific performance and broad adaptability without requiring full parallelization for all cases.
Data Source
AI summary
A first group of elements is element-wise multiplied with a second group of elements using a plurality of multipliers belonging to a matrix multiplication hardware unit. Results of the plurality of multipliers are added together using a hierarchical tree of adders belonging to the matrix multiplication hardware unit and a final result of the hierarchical tree of adders or any of a plurality of intermediate results of the hierarchical tree of adders is selectively provided for use in determining an output result matrix. A control unit is used to instruct the matrix multiplication hardware unit to perform a plurality of different matrix multiplications in parallel by using a combined matrix that includes elements of a plurality of different operand matrices and utilize one or more selected ones of the intermediate results of the hierarchical tree of adders for use in determining the output result matrix that includes different groups of elements representing different multiplication results corresponding to different ones of the different operand matrices.


