Shared SFU Architecture for AI Chip Area Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI processors face inefficiencies in executing complex computations due to the time-consuming and labor-intensive process of combining basic arithmetic and logical operations, which also occupies large chip areas and increases costs.
Innovation Solution
A complex computing device with a shared Special Function Unit (SFU) architecture, where multiple AI processor cores share the same data paths for complex computing instructions, reducing the area and power consumption of the AI chip.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If basic arithmetic and logical operations are combined to implement complex computations in AI processors, then complex operations can be performed, but the execution process becomes time-consuming and labor-intensive
Solution Approach 1:
The patent merges multiple basic arithmetic and logical operations into a single complex computing instruction. The computing device receives a complex computing instruction that encompasses multiple operations (e.g., multiplication, addition, activation functions) and executes them as one unified operation, thereby improving execution efficiency while maintaining complex computation capability
Solution Approach 2:
The computing device is designed with universal computing components that can handle various types of complex operations through a single instruction interface. The device performs different complex computations (convolution, matrix multiplication, activation functions) using the same hardware infrastructure, enabling multi-functionality without sacrificing performance
2Productivity
If dedicated complex computing units are provided for each processor core, then complex operations can be executed efficiently, but the chip area occupied increases and costs rise
Solution Approach 1:
The patent merges the complex computing units from multiple processor cores into a shared resource. Instead of each core having its own dedicated complex computing unit, the system implements a shared computing device that multiple cores can access, thereby reducing the total chip area while maintaining execution efficiency through coordinated access and arbitration
Solution Approach 2:
The shared computing device is designed with universal functionality to serve multiple processor cores. It can handle complex computing tasks from any of the N processor cores through a unified interface and arbitration mechanism, providing multi-functional capability that eliminates the need for duplicate dedicated units in each core
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present application discloses a complex computing device, a complex computing method, an artificial intelligence chip and an electronic apparatus, and relates to a field of artificial intelligence chips. One of the solutions includes: an input interface receives complex computing instructions and arbitrates each complex computing instruction to a corresponding computing component respectively, according to the computing types in the respective complex computing instructions; each computing component is connected to the input interface, acquires a source operand from a complex computing instruction to perform complex computing, and generates computing result instruction to feed back to an output interface; the output interface arbitrates the computing result in each computing result instruction to the corresponding instruction source respectively, according to the instruction source identifier in each computing result instruction.