Shared SFU Architecture for AI Chip Area Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI processors face inefficiencies in executing complex computations due to the time-consuming and labor-intensive process of combining basic arithmetic and logical operations, which also occupies large chip areas and increases costs.

Innovation Solution

A complex computing device with a shared Special Function Unit (SFU) architecture, where multiple AI processor cores share the same data paths for complex computing instructions, reducing the area and power consumption of the AI chip.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If basic arithmetic and logical operations are combined to implement complex computations in AI processors, then complex operations can be performed, but the execution process becomes time-consuming and labor-intensive

Engineering Contradiction:
Improvecomplex computation capabilityVSAvoidexecution efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent merges multiple basic arithmetic and logical operations into a single complex computing instruction. The computing device receives a complex computing instruction that encompasses multiple operations (e.g., multiplication, addition, activation functions) and executes them as one unified operation, thereby improving execution efficiency while maintaining complex computation capability

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The computing device is designed with universal computing components that can handle various types of complex operations through a single instruction interface. The device performs different complex computations (convolution, matrix multiplication, activation functions) using the same hardware infrastructure, enabling multi-functionality without sacrificing performance

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If dedicated complex computing units are provided for each processor core, then complex operations can be executed efficiently, but the chip area occupied increases and costs rise

Engineering Contradiction:
Improveexecution efficiencyVSAvoidchip area
Core Design Contradiction:
ProductivityVSArea of stationary object

Solution Approach 1:

The patent merges the complex computing units from multiple processor cores into a shared resource. Instead of each core having its own dedicated complex computing unit, the system implements a shared computing device that multiple cores can access, thereby reducing the total chip area while maintaining execution efficiency through coordinated access and arbitration

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The shared computing device is designed with universal functionality to serve multiple processor cores. It can handle complex computing tasks from any of the N processor cores through a unified interface and arbitration mechanism, providing multi-functional capability that eliminates the need for duplicate dedicated units in each core

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3933586B1Complex computing device, complex computing method, artificial intelligence chip and electronic apparatus
Publication Date: 2025.03.05 KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD
  • EP3933586B1 patent drawingFigure 1
  • EP3933586B1 patent drawingFigure 2
  • EP3933586B1 patent drawingFigure 3

AI summary

The present application discloses a complex computing device, a complex computing method, an artificial intelligence chip and an electronic apparatus, and relates to a field of artificial intelligence chips. One of the solutions includes: an input interface receives complex computing instructions and arbitrates each complex computing instruction to a corresponding computing component respectively, according to the computing types in the respective complex computing instructions; each computing component is connected to the input interface, acquires a source operand from a complex computing instruction to perform complex computing, and generates computing result instruction to feed back to an output interface; the output interface arbitrates the computing result in each computing result instruction to the corresponding instruction source respectively, according to the instruction source identifier in each computing result instruction.