Vector Inner Product Computing Apparatus for Machine Learning Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies for performing floating-point number vector inner product computations are inefficient and require excessive hardware, limiting their speed and scalability in applications such as machine learning algorithms.
Innovation Solution
A computing apparatus and method that includes a multiplication unit with floating-point multipliers and an addition unit, featuring an update mechanism with a second adder and register to efficiently perform vector inner product computations across various floating-point data formats, reducing hardware requirements and supporting different data formats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If existing multiplication and addition apparatuses are used for floating-point number vector inner product computation, then the computation can be performed, but the execution efficiency is insufficient and hardware resources are excessive
Solution Approach 1:
The patent divides the vector inner product computation into multiple stages: a multiplication unit that computes partial products of vector elements, and an addition unit that accumulates these partial products. This segmentation allows parallel computation of multiple partial products simultaneously, improving execution efficiency while using a manageable number of hardware resources.
Solution Approach 2:
The patent introduces a hierarchical structure where the addition unit includes both a first adder for initial accumulation and a second adder for final summation, operating at different computational stages. This multi-dimensional approach to addition enables efficient handling of intermediate results and achieves high-speed computation with reduced hardware complexity.
2Speed
If more multiplication and addition apparatuses are arranged to improve computation speed, then execution efficiency improves, but hardware arrangement area increases excessively
Solution Approach 1:
The patent segments the computation into parallel multiplication operations followed by hierarchical addition. The multiplication unit can process multiple element pairs simultaneously, and the addition unit efficiently accumulates results in stages, achieving high computation speed without requiring a large number of adders that would increase hardware area.
Solution Approach 2:
The patent ensures continuous computation by having the multiplication unit continuously generate partial products that are immediately fed to the addition unit for accumulation. This continuous pipeline operation maintains high computation speed while using a compact, efficient hardware structure that minimizes arrangement area.
Data Source
AI summary
The present disclosure relates to a computing apparatus, a method and an integrated circuit chip for a vector inner product, where the computing apparatus may be included in a combined processing apparatus. The combined processing apparatus may further include a general interconnection interface and other processing apparatus. The computing apparatus interacts with other processing apparatus to jointly complete a computing operation specified by a user. The combined processing apparatus may further include a storage apparatus, where the storage apparatus is respectively connected to the computing apparatus and other processing apparatus, and the storage apparatus is used for storing data of the computing apparatus and other processing apparatus.


